From Manual to Minutes: 4 Decisions That Built a Production AI Order System in 90 Days
We automated a workflow that was costing us hours every day. It took 90 days. Here are the four decisions that made or broke the project.
We automated a workflow that was costing us hours every day. It took 90 days. Here's what actually mattered:
The problem: every order arrived as an email with attachments — PDFs, spreadsheets, Word docs. Someone had to read them, look up IDs, check for errors, and manually enter data into our system.
Precise. Repetitive. Error-prone. Expensive.
So we built an AI system to handle it end-to-end.
Four decisions made or broke the project:
Decision 1: Test the hard part first.
Before building any infrastructure, we pointed an AI at a real attachment and asked it to extract structured data. It worked — imperfectly, but well enough to prove the core was possible.
Most teams build the scaffolding before validating the foundation. We validated first. Saved us months.
Decision 2: Build an evaluator before building features.
Week 6: we built an agent that compared AI output against known-good orders and measured accuracy.
This was the most valuable thing we built. Every improvement we made after that was measurable. Without it, we'd have been optimizing by gut feel.
You can't improve what you don't measure. Build the feedback loop early.
Decision 3: Keep knowledge out of code.
This is the one most teams skip.
Business rules, domain vocabulary, customer-specific formats — we put all of it in Markdown files, not code. When a rule changes, someone edits a file. No deployment. No risk of breaking something else.
The AI reads Markdown as fluently as it reads Python. Use that.
Decision 4: Local data over live APIs.
We sync all reference data to disk on a schedule. At processing time, the AI queries local files — no live API calls.
This had three effects we didn't fully anticipate:
- Evaluation became fast (batch of 20 test cases run parallel in minutes, not hours)
- The AI could fall back to raw data when standard lookups failed — and nearly always recovered
- Success rate went from ~90% to nearly 100%
That last point is the one worth sitting with. The difference between 90% and 100% isn't a better model. It's giving the model the right data and the right fallback strategy.
The mechanism: we put Claude Code in the driver's seat. When a standard lookup fails, Claude Code doesn't give up — it follows a defined escalation path: grep the raw data files, scan order history for similar cases, re-read the original input. It's not improvising; it's following auditable instructions with direct file access. That combination is what closed the gap.
The result:
Processing time: hours → minutes.
Human review time: minutes → seconds (for clean orders).
Accuracy: continuously improving through an automated loop.
No new code required to onboard a new customer — just a new knowledge file.
What I'd tell someone starting this today:
Build the evaluator in week one, not month two. Design knowledge separation into your first architecture, not your second. Don't confuse "the AI is smart" with "the system is reliable" — reliability comes from structure, not capability.
The AI is a powerful primitive. The system design is the hard part.
We built this in 90 days with a small team, iterating fast with AI-assisted development. The same principles apply to any workflow where humans are processing structured information from unstructured inputs: invoices, support tickets, contracts, forms.
What's the workflow in your organization that's still manual because "it's too complex to automate"? I'd bet the architecture above handles it.
Originally published on LinkedIn on March 1, 2026.