Freelance & Fractional
AI Receptionist &
Voice Agent Development
I design, build, and rescue production voice AI agents — the kind that answer real phone calls, book real appointments, and get measured on real dollars. This is what I do full-time at a YC-backed voice AI company. I take on a small number of outside projects.
No agency overhead. You work directly with the engineer.
Where the experience comes from
- Engineering Manager at Avoca (YC '23) — I lead the Deployed Engineering team shipping AI voice agents for home-services companies. Voice agents in production are my day job, not a side interest.
- Staff Engineer at Twilio — built on Twilio Verify. Telephony, carrier behaviour, and what actually breaks on a real phone network.
- PayPal risk platform — infrastructure handling 200M+ payments a day. Systems that cannot go down.
- Zalando, SingleStore, Wayfair, Nutanix, HP — near real-time ML for fraud detection, control plane and billing, supply chain systems at scale.
- Founder of QRCodeStack — I ship and operate my own product, so I understand cost per unit and churn, not just architecture diagrams.
What an AI receptionist actually needs to work
Most AI receptionist projects fail for the same four reasons, and none of them are the model. If you are evaluating building one, these are the things to get right first.
- It has to answer fast enough to feel human. Past roughly 800ms to first word, callers start talking over the agent and the conversation falls apart. That is an architecture decision, not a tuning one.
- It has to know when you stopped talking. Bad endpointing cuts people off mid-sentence. This single issue causes more abandoned calls than any wrong answer.
- It has to actually write to your calendar. An AI phone answering service that says "you are booked" without the booking existing is worse than no agent at all. Verify downstream, in the system of record.
- It has to know when to hand off. Warm transfer with context, not a cold dump into a queue after the caller has already explained themselves twice.
What I build
-
01
AI receptionist & phone answering agent — zero to production
Full build of a working AI receptionist: telephony and SIP, speech-to-text, LLM orchestration, text-to-speech, tool and function calls into your systems, warm handoff to a human, and the dashboard your ops team will actually use. Delivered as code you own, not a locked platform. -
02
Evals, testing & observability for voice agents
The thing almost everyone skips. A regression suite of real call transcripts, automated scoring for intent coverage, task success and hallucination, plus per-call tracing — so you can change a prompt or swap a model on a Tuesday and know by Wednesday whether it got worse. I have written the field guide on this. -
03
Voice agent audit & rescue
You already have an agent and it is underperforming. I run a structured diagnostic across latency, endpointing, containment, tool-call reliability, and cost per resolved contact, then hand you a prioritised fix list with the estimated impact of each item. -
04
AI agent architecture & platform selection
Should you build on Vapi, Retell, LiveKit, Pipecat, or roll your own orchestration? Single agent, or a graph of specialised ones? I have run these architectures side by side and will give you a straight recommendation with the trade-offs written down. -
05
Fractional AI engineering leadership
Part-time technical leadership for a team already building agents — architecture reviews, hiring and interview loops, delivery process, and mentoring your engineers so the capability stays after I leave. -
06
Backend & data infrastructure behind the agent
Voice agents are only as good as what they can reach. PostgreSQL, Kafka, event pipelines, queueing and retries, CRM and scheduling integrations — the unglamorous layer that decides whether the agent can actually complete the task.
Voice AI for your industry
The hard parts change a lot by vertical. These are the ones where I have the most context.
How much does AI voice agent development cost?
Straight answer: it depends on four things, and any quote you get without them being discussed is a guess.
-
1
Intent surface
One job — "book an appointment" — is a fixed-scope project. Twelve intents with branching logic is a different animal entirely. -
2
Number of systems it must touch
An agent that only talks is cheap. An agent that reads your calendar, writes to your CRM, and takes a payment is where the real work lives. -
3
Answering vs. acting
A wrong answer is embarrassing. A wrong action — a cancelled appointment, a double charge — is expensive. Acting agents need far more eval and guardrail work. -
4
Compliance bar
Call recording consent, PII handling, retention, and audit trails. In healthcare and finance this is a meaningful share of total build time, not a checkbox at the end.
How engagements run
-
01
Free 30-minute scoping call
What you are building, where it is stuck, what "working" means in numbers. If I am not the right person you will hear that on this call. -
02
Written scope and price
A short document: what gets built, what does not, what I need from you, how we know it worked, and the cost. Fixed price wherever the scope allows it. -
03
Build in weekly increments
Something demoable every week and a written update. No month-long silences ending in a surprise. -
04
Handover you can actually maintain
Code you own, the eval harness, a runbook, and a walkthrough session with your engineers. The goal is that you do not need me on retainer forever.
Read my work before you hire me
The fastest way to judge whether I know this domain is to read what I have written about it.
- The Essential Metrics for Voice AI Agents — From Milliseconds to Dollars — the five-layer scorecard I use to judge whether an agent is actually working.
- Evals, Testing, and Observability for Voice AI Agents: A Field Guide — how to stop shipping regressions you cannot see.
- What Are Evals? Illustrated by a Real Voice AI Agent — the plain-English version for non-engineers on your team.
- What Is an AI Agent Harness? — plus one built in 80 lines of Python.
- Agent Loops, Swarms, and Graphs — the architecture comparison, run rather than theorised.
- The Anatomy of a Great Prompt — why most agent quality problems are prompt problems.
Frequently asked questions
See the four cost drivers above. A single-intent agent on an existing platform is a small fixed-scope project; a multi-intent agent that writes into your CRM, takes payments, and needs an audit trail is a multi-month build. I scope against those axes before quoting rather than pricing off a template.
All of them, depending on what you need. Platforms like Vapi and Retell get you to a working agent fastest and are the right call for most first builds. A custom orchestration layer on LiveKit or Pipecat earns its keep when you need control over turn-taking, interruption handling, or per-call cost at volume. I will tell you honestly which side of that line you are on — including when the answer is that you do not need me.
Almost never the model. It is latency above roughly 800ms to first token, endpointing that cuts callers off mid-sentence, silent tool-call failures that make the agent confidently wrong, and no eval harness — so nobody notices quality dropped until customers complain. Most rescue work is fixing those four things, not swapping the LLM.
Yes, and it is usually the better outcome. My day job is leading a Deployed Engineering team, so embedding with your engineers, setting up architecture and eval process, then handing over is a normal engagement shape rather than a special request.
Teams in the US, UK, Europe, and India, with overlap hours agreed up front. Engagements are remote and mostly asynchronous, with one fixed weekly sync.
Email [email protected] with what you are building and where it is stuck. First call is a free 30-minute scoping conversation. If it is not a fit, I will say so and point you somewhere better.
Have a voice AI project that needs to actually ship?
Email [email protected]Or find me on LinkedIn. I reply to everything that is not a template.