I find what's wrong, fix it, and keep it running.
The model is rarely the issue — the seams between systems are. This is what I get called in to fix.
Invents an answer the moment it hits something it doesn't know.
Ask "is it delayed?" and it no longer knows what you meant two messages ago.
An API fails and the agent improvises around it instead of stopping.
Out-of-scope cases get a confident wrong answer instead of a person.
One retry loop or one abuser burns through the token budget.
State doesn't survive a redeploy, so every update makes it a little worse.
Travelers talk to it to check a flight, find a gate, or sort a missed connection — in English and Turkish, on live flight data. Guardrails keep it from inventing a flight number it doesn't have.
Built for a client handling hundreds of inbound messages a day. Intent classification splits real questions from noise, long-term memory remembers every customer, and low-confidence cases escalate to a human.
A Spanish-language support bot that routes customer DMs into one Crisp inbox, with two-way sync so replies post back to Telegram.
A peer-to-peer escrow protocol on BSC Mainnet: ~2,000 lines of Solidity, 68/68 tests passing, and an independent audit with all 8 findings resolved before launch.
I find the real failure modes before quoting a fix.
Clear price, clear timeline. If I haven't shipped it before, I say so.
Designed around the failure modes from day one.
Documented code and a runbook your team can run without me.
Then the audit says so, and you keep the findings. Sometimes the right answer is a different architecture.
An audit usually turns around in days. Fixes are scoped from there.
Either. Small fixes work best fixed-price; larger work runs hourly or in milestones.
Claude and GPT APIs, RAG, MCP, multi-agent orchestration, voice agents, Python and FastAPI, Next.js, vector databases, and the webhook and integration plumbing that ties it together.
Yes. The reliability work is the same either way.
I ship AI that runs in production, handling real customers and real transactions. You get documented code, a runbook, and honest scoping.
Tell me what's going wrong. I'll tell you what it'll take to fix, and whether it's worth it.
Usually replies within a day · Taipei (GMT+8)