01 — Practice
AI products that are measured, not demonstrated.
A convincing demo and a dependable product are different artifacts. We build the second kind: retrieval that cites its sources, agents whose permissions are explicit, and an evaluation suite that has to pass before anything ships.
Approach
How an engagement actually runs.
Scope against a baseline
Before any model work, we establish how the task is done today and what "correct" means numerically. Without that number, every later claim of improvement is a guess.
Retrieval before fine-tuning
Most failures blamed on the model are retrieval failures. We build the corpus pipeline — chunking, embedding, ranking, citation — and measure it on its own.
Bounded agents
When a system takes actions, each tool gets an explicit permission scope and an audit record. Irreversible operations require confirmation by construction, not by prompt.
Evaluation as a gate
A regression suite runs on every change. Prompt edits are code changes and get reviewed as such. Nothing reaches production on the strength of a demo.
What we build with
The layers we own in a delivery.
- Model access
- One routing layer over the providers you approve, swappable without a rewrite
- Retrieval
- Vector and hybrid search over your own corpus, with citations back to source
- Orchestration
- Tool-calling agents with typed interfaces and bounded permissions
- Evaluation
- Golden sets, LLM-as-judge with human calibration, regression gating in CI
- Interface
- The part people actually touch — designed, not bolted on at the end
- Observability
- Traced calls, token accounting, latency and failure budgets
The routing layer is deliberate. Model providers leapfrog each other every few months, and a system welded to one of them is a system that gets worse by standing still. Yours should be able to change vendors without a rewrite — and you should know which one is answering.
Have a task you think a model could take over?
Send us the workflow and what happens when it goes wrong. If the honest answer is that it is not an AI problem, we will say that instead of building something.