Philipp Guldimann

Philipp Guldimann

Machine Learning Engineer · Zürich, Switzerland

Open to AI engineering roles — Zürich or remote

I build the evaluation and data infrastructure that makes AI systems survive production.

I am joint first author of COMPL-AI, the first technical interpretation of the EU AI Act and an open benchmarking suite built on it. At LatticeFlow AI I built LLM evaluation infrastructure — 23 evaluators across 10+ models — that cut QA cycles from two days to four hours. At Omnilex I built multi-jurisdiction legal AI pipelines processing roughly one million documents across three countries. I hold an MSc in Machine Intelligence from ETH Zürich.

  • LLM evaluation
  • Data pipelines
  • Trustworthy AI
  • Python / TypeScript
  • AWS / Azure

Writing

I write about the layer around a model in production, and where it breaks.

  • Running the tube backwards

    · Speech from first principles · Part 2 · 16 min read

    The acoustic model of a speech recogniser, derived as the inverse of the one that makes the sound.

  • How a tube becomes a vowel

    · Speech from first principles · Part 1 · 20 min read

    Deriving the acoustic model of the human voice from the wave equation — with every step you can hear and take apart.

  • Your tool exited 0. It did nothing.

    · Harness engineering · Part 1 · 8 min read

    What a misbehaving RGB controller taught me about designing tools for language models.

Experience

Data Engineer — Omnilex

  • Built ingestion and transformation pipelines for legal content across 3 jurisdictions (Switzerland, Germany, Austria), processing ~1M documents from APIs, scraping outputs, and bulk sources.
  • Developed TypeScript-based data workflows for normalization, citation-aware chunking, embeddings, classification, and entity extraction.
  • Contributed to RAG-ready indexing and Azure-based search infrastructure for precise, traceable legal AI responses.
  • Introduced data contracts and validation checks that caught 50K+ duplicate entries before they reached production.

Machine Learning Engineer — LatticeFlow AI

  • Built v0 evaluation infrastructure for a new AI product from scratch: 23 evaluators, integrations with multiple data and chat-model providers, assessing 10+ LLMs.
  • Automated evaluation workflows for QA and regression testing, reducing review cycles from ~2 days of manual inspection to ~4 hours.
  • Extended evaluators with targeted dataset generation to increase coverage across model behaviours and failure modes.

Research

COMPL-AI — Benchmarking LLM Compliance with the EU AI Act

Joint first author (equal contribution) and lead author on COMPL-AI — the first technical interpretation of the EU AI Act, mapping its six ethical principles onto 27 concrete benchmarks and evaluating 12 prominent LLMs against them. Worked on benchmark design, evaluation pipelines, and model integration via Hugging Face Transformers. The framework is positioned as a reference point for the EU’s GPAI Code of Practice.

Read the paper (arXiv:2410.07959) · Code on GitHub

Projects

DomainForge — Relational domain-model induction

Give a system a corpus from a company it knows nothing about, and no schema. Agents propose changes to a shared knowledge state; a minimum-description-length energy and a Metropolis-Hastings rule decide what survives, so the ontology is annealed rather than fixed in advance. Thirteen reversible operators, each with its inverse asserted by round-trip tests.

Project page · Interactive walkthrough

Evidara — Legal document intelligence platform

Search over legal sources that differ in authority, jurisdiction and trust, built around one decision: canonical truth is separated from the serving projection, so the index is a rebuildable view rather than the system of record. Thirteen services, moved from Google Cloud onto self-hosted Kubernetes, with 59 architecture decision records and two dozen automated checkers holding the boundaries in place.

Project page

Education

MSc Computer Science — Machine Intelligence, ETH Zürich

Thesis (top grade): Speech Recognition for Children with Congenital Disorders Using Adaptive Methods — adapting Whisper to non-normative child speech from a single speaker.

Read the case study →

BSc Computer Science, ETH Zürich

Thesis (top grade): Detecting Disinformation on Twitter Targeting Non-Profit Organisations, in collaboration with the ICRC. Thesis (PDF)

Technical skills

   
Languages Python, TypeScript / JavaScript, SQL
LLM / AI LLM evaluation, RAG, embeddings, Hugging Face Transformers, MCP, PyTorch, LoRA / PEFT
Data & orchestration Dagster, PostgreSQL, pgvector, OpenSearch, Azure AI Search, Delta Lake, NATS
Backend & cloud FastAPI, NestJS, Next.js, Node.js, Docker, Kubernetes, Terraform, Argo CD, AWS, Azure, CI/CD

Earlier projects

Computational Intelligence Lab — Text Classification

Report (PDF)

Contact

Email is the fastest way to reach me.

phil.guldimann@gmail.com LinkedIn GitHub