Built On Our Own Production Systems
Site Factory's content pipeline is agent-driven automation at real scale — we build for ourselves before we sell it.
Software that takes the action, not just describes it.
Custom AI agents built for a specific job — tool-use, memory, and a real evaluation set to prove they work — with a human handoff for the cases the model shouldn't handle alone. Delivered as code you own.
Most "AI agent" projects fail for the same reason: someone wires a model to a prompt, it works in the demo, and it falls apart on the tenth real case that doesn't look like the first nine. An agent that's actually useful in production needs three things a demo skips — tool-use that's tested against real systems, memory that persists correctly across a multi-step task, and an honest evaluation of where it breaks.
The economical version isn't "automate everything." It's identifying the repetitive 80% of a process an agent can genuinely own, and leaving the judgment-heavy 20% with a person — then building the handoff between them properly instead of pretending the agent can do it all.
A prompt that impresses in a demo often fails on the messy, real-world version of the same task — evaluation is what catches that before clients do.
Knowing exactly when to route to a human, and doing it cleanly, is what separates a reliable agent from an unreliable one.
Building around one provider means rebuilding when a better or cheaper model appears. Multi-provider routing avoids that from day one.
Four reasons the agent still works six months after launch, not just at the demo.
Site Factory's content pipeline is agent-driven automation at real scale — we build for ourselves before we sell it.
A real test set of cases with known correct answers, not a demo run a handful of times and called done.
Explicit handoff to a human for the cases the model shouldn't be trusted with alone — designed in, not bolted on after a failure.
Multi-provider routing from day one, so switching models later doesn't mean a rebuild.
An agent built and tested against the real version of the task, not a polished demo.
The task broken into what an agent can own reliably and what still needs a person, before any code is written.
Connections to the APIs, databases, or internal systems the agent needs to actually take action, not just talk about it.
State that persists correctly across a multi-step task, so the agent doesn't lose track halfway through.
Real cases with known correct answers, tested before launch and reused to catch regressions later.
Explicit rules for when the agent hands off to a human instead of guessing on a case it can't handle.
Built to switch models without a rebuild, so you're never stuck with one vendor's pricing or performance.
Visibility into what the agent is doing in production, not a black box you have to trust blindly.
Source code, prompts, and infrastructure configuration delivered with documentation — no licensed layer.
Two to eight weeks depending on tool complexity and how much evaluation the task requires.
Confirm what the agent should own versus what stays with a person, based on the real process, not an assumption.
Week 1The specific APIs and systems the agent needs to take real action, scoped and connected.
Week 1-3Agent logic built alongside a real test set of cases with known correct outcomes.
Week 2-5Fallback rules defined and tested against edge cases before anything reaches production.
Week 4-6Deployed with monitoring in place, and full documentation delivered to your team.
Week 6-8For teams with a repetitive process that's too irregular for a simple script but too high-volume for a person.
High-volume, repetitive enquiries that need a real answer, with escalation to a human for anything unusual.
Manual data entry, reconciliation, or reporting work that's consumed the same expensive hours for years.
Teams drowning in inbound volume who need the repetitive triage handled before a person gets involved.
Products that want an agent as part of the offering itself, not just an internal efficiency tool.
Teams whose delivery model doesn't scale linearly with headcount without automating part of the process.
Teams who've seen an AI prototype impress in a meeting and then never make it to production.
Scoped on tool complexity and evaluation depth — a fixed project, not an open-ended retainer.
$2,500-5,000one-off build
One well-defined task with a small number of tool integrations.
$6,000-15,000one-off build
A process chained across several tools and decision points.
$500-2,000per month, optional
Keeping the agent tuned as your process and available models evolve.
Exact pricing depends on tool count and how rigorous the evaluation needs to be — confirmed after a scoping call.
Describe the task. We'll tell you honestly whether it's economical to automate, and roughly what the build would involve.
Evaluated before launch, with guardrails where the model shouldn't be trusted alone.
Software that can take an action, not just generate text — calling APIs, updating records, or making a decision within defined limits, with a human step where the model shouldn't be trusted alone.
A chatbot answers questions. An agent does work: it can look something up, take an action, and report back, chained across multiple steps without a person driving each one.
Whichever fits the task, with multi-provider routing built in from the start — you are not locked into one model vendor if a better or cheaper option appears later.
We build an evaluation set during the project — real cases with known correct answers — and test against it before launch, plus guardrails and fallbacks for edge cases the model handles badly.
It hands off to a human step rather than guessing. Deciding where that line sits is part of the design work, not an afterthought.
Both, where needed — the reasoning layer and the specific tool integrations (APIs, databases, internal systems) it needs to actually do the job.
You do — agent logic, tool integrations, and infrastructure configuration, delivered with documentation. No licensed layer you keep paying us for.
Model and API usage billed to your own provider accounts, plus hosting. Sized during scoping so there are no surprises once it's live.
Almost always part of one. The economical agents take over the repetitive 80% of a process and leave the judgment-heavy 20% with a person — trying to automate everything usually costs more than it saves.
Two to eight weeks depending on how many tools the agent needs to use and how much evaluation the task calls for before it can be trusted in production.