AI engineering leader | Building production AI agents and ML systems | UBS Director, AI Innovation | GCP | Python
No open PRs to review
A collection of notebooks/recipes showcasing some fun and effective ways of using Claude.
Companion code for: Beyond Hallucination — A System-Level Failure Taxonomy for Production LLMs
🙌 OpenHands: AI-Driven Development
Build resilient language agents as graphs.
:zap: Dynamically generated stats for your github readmes
Vertex AI utility scripts — model serving, BigQuery ML integration, Cloud Run templates
Document classification and extraction for financial documents with interpretability outputs
Production RAG: hierarchical chunking, query rewriting, re-ranking, uncertainty thresholding
GitHub Actions + Vertex AI ML pipelines — shadow deploys, rollback, audit trail
Production LLM evaluation: golden datasets, model-as-judge scoring, CI/CD regression gates
- feat(evals): add model-as-judge evaluation pipeline with CI/CD regression gates
Hi @nikblanchet and @cj-ant — flagging this for your attention when you get a chance. This adds a second cookbook to the evals/ directory: a practical model-as-judge evaluation pipeline for teams shi…
- feat(agent-sdk): add compliance-aware agent with human-in-the-loop for regulated environments
Hi @nikblanchet and @cj-ant — wanted to flag this for your attention when you get a chance. This adds notebook #08 to the Agent SDK series: a compliance-aware multi-agent system with human-in-the-loo…