AI Reliability · Independent reviews of LLM products · Software Architecture & Development
I find where your AI product breaks, and the design decisions that let it break.
Independent reviews for teams whose LLM feature is live but unmeasured. I reproduce failures from your real outputs, review the architecture around the model, and trace each one to the design decision behind it. Anything I can't verify stays out of the report.
The AI Reliability Review
Fixed scope, two weeks.
+1–3 days if you don't have tracing yet.
01 · You provide
Three inputs:
- Repo access
- A slice of real outputs — or a way to generate them
- 2–3 hours of team time
02 · I review
Both sides of the system:
- Behavior — failures in your real outputs, reproduced and counted
- Design — the architecture around the model: data flow, retrieval, prompts, fallbacks
- The trace — each failure tied to the design decision behind it
03 · You leave with
Five deliverables:
- Failure inventory
- Architecture findings
- Ranked fix roadmap
- Starter regression set
- Executive readout
The method
How AI fits in the work, and why the findings hold.
I use AI heavily in every review: it reads more outputs, tries more angles, and covers more of the system than I could alone. But the judgment doesn't get outsourced, and neither does the verification. Every finding is reproduced by hand before it reaches the report.
- Search
AI-accelerated analysis
AI reads more of your outputs and your codebase than a human reviewer could. It surfaces candidates: possible failures, suspect paths, patterns worth a look.
- Verify
Candidates become findings
AI-assisted analysis produces findings that are plausible and wrong. So a candidate becomes a finding only when I reproduce it and trace it to a cause. What I can't verify doesn't ship.
- Evidence
Findings you can re-run
Each finding carries its trace and its reproduction steps. Your team can check my work without me in the room.
- Design
Down to the decision
Failures get traced to the design decision that allowed them: retrieval, data flow, fallbacks, orchestration. The fix plan changes the system, and the regression set keeps it honest.
Selected work
The work behind the reviews: where the architectural and engineering depth comes from. Read more case studies
- 2026
Verification-first AI audit tool
Architect & Author · Open source
An open-source audit skill that reviews a codebase's auth layer for vendor lock-in risk. AI does the reading; every finding is quoted from the code and verified against live runs, and the harness that tests the auditor is mutation-tested. The method behind my reviews, working in public.
- 2024—25
Backend for an AI-agents platform
Lead Backend Architect · Bullseye Web3 Studio
Architected the event-driven microservices backend behind two greenfield products, including A1X, a platform where users create and run their own AI agents. Zero to 150,000+ registered users; GCP stayed under $500/month across 10+ services, mostly by deciding what not to build before product-market fit.
- 2020—25
LLM chatbot in a health product
Fractional CTO · Parentool
Shipped a production LLM chatbot (OpenAI, structured outputs) inside a health-tech product I ran end to end: 10,000+ users, 7% paid conversion, peak at #3 in App Store Health & Fitness. A domain where wrong answers carry real cost.
- 2025—26
Auth system re-architecture
Lead Architect & Engineer · Pie Insurance
Owned the authentication track of a unified frontend re-architecture across a Partner Portal of 100+ backend microservices. Wrote the ADRs (framework, token storage, OAuth, multi-pool Cognito) and migrated the legacy Amplify/SRP auth to a modern OAuth flow on Cognito Managed Login.
What teams say
From the people who hired me: enterprise leads, founders, clients.
Before diving into the code, he takes the time to thoroughly understand the business requirements — that meticulous upfront analysis lets him anticipate complex edge cases and architectural roadblocks long before they reach production. He has the rare maturity to provide constructive pushback when necessary.
Shilpi ReddyEngineering Leader · Pie Insurance He helped me transform my initial concepts into clear, structured documentation and provided insights that added real value to the project. What impressed me most was his ability to communicate technical concepts in a way that was accessible to everyone, ensuring alignment across the board. A consummate professional.
Emin EskiocakPropTech Entrepreneur He delivered a highly functional, almost bugless solution in the exact timeline we agreed — and could explain to us, non-technical people, everything happening in the backend.
Petruța CosteaFounder · Parentool
Writing
Notes from production work, written up in full. Read the blog
Book a call
Ready to find out where your AI product breaks?
Pick your situation, pick a time. The intro call is 30 minutes.
01 · Where are you with your AI feature?
02 · Pick a time
Intro call
30 min · Google Meet