Philipp Guldimann

AI Harness Engineer · Zürich, Switzerland

I build the scaffolding that makes AI systems survive production — evaluation, agent reliability, and the data pipelines underneath.

At a Zürich AI startup I built LLM evaluation infrastructure — 23 evaluators across 10+ models — that cut QA cycles from two days to four hours. I’m now building multi-jurisdiction legal AI pipelines processing roughly one million documents across three countries. My roots are in trustworthy AI research: I was a core contributor to COMPL-AI, an open benchmark for LLM compliance with the EU AI Act. I hold an MSc in Machine Intelligence from ETH Zürich, and I care about pragmatic, scalable systems that ship and prove their value with metrics.

  • LLM evaluation
  • Agent reliability
  • Data pipelines
  • Evals & MLOps
  • Platform engineering
  • Trustworthy AI
  • Python / TypeScript
  • AWS / Azure

Writing

I write about AI harness engineering — agent loops, tool schemas, context management, evals — and where they break in practice.

Experience

Data Engineer — Omnilex

  • Built ingestion and transformation pipelines for legal content across 3 jurisdictions (Switzerland, Germany, Austria), processing ~1M documents from APIs, scraping outputs, and bulk sources.
  • Developed TypeScript-based data workflows for normalization, citation-aware chunking, embeddings, classification, and entity extraction.
  • Contributed to RAG-ready indexing and Azure-based search infrastructure for precise, traceable legal AI responses.
  • Introduced data contracts and validation checks that caught 50K+ duplicate entries before they reached production.

Machine Learning Engineer — LatticeFlow AI

  • Built v0 evaluation infrastructure for a new AI product from scratch: 23 evaluators, integrations with multiple data and chat-model providers, assessing 10+ LLMs.
  • Automated evaluation workflows for QA and regression testing, reducing review cycles from ~2 days of manual inspection to ~4 hours.
  • Extended evaluators with targeted dataset generation to increase coverage across model behaviours and failure modes.

Research

COMPL-AI — Benchmarking LLM Compliance with the EU AI Act

Core contributor to COMPL-AI, evaluating 10+ models across 20 benchmarks spanning capabilities, cybersecurity, privacy, and bias/fairness. Worked on benchmark design, evaluation pipelines, and model integration via Hugging Face Transformers.

Read the paper (arXiv:2410.07959) · Code on GitHub

Education

MSc Computer Science — Machine Intelligence, ETH Zürich

Thesis (top grade): Speech Recognition for Children with Congenital Disorders Using Adaptive Methods — adapting Whisper to non-normative child speech from a single speaker.

Read the case study →

BSc Computer Science, ETH Zürich

Thesis (top grade): Detecting Disinformation on Twitter Targeting Non-Profit Organisations, in collaboration with the ICRC. Thesis (PDF)

Technical skills

   
Languages Python, TypeScript / JavaScript, SQL
LLM / AI LLM evaluation, RAG, embeddings, Hugging Face Transformers, LangChain / LangGraph, MCP, PyTorch
Data & orchestration Airflow, Dagster, vector databases (FAISS, Weaviate, Pinecone), PostgreSQL
Backend & cloud FastAPI, Node.js, Docker, Kubernetes, AWS, Azure, CI/CD

Earlier projects

Computational Intelligence Lab — Text Classification

Report (PDF)