Skip to content
Engineering Square
AI in Production

We run AI agents on our own operations. Here is what that takes.

What it actually takes to run AI agents in production on real operations — lessons from operating a work-orchestration MCP server, an agentic HR platform, and a self-evolving software platform inside our own company.

Published September 30, 2026Read 5 minBy Engineering Square

Running AI agents in production on real operations takes four things no launch post prepares you for: a written definition of what each agent may read, write, and decide; an approval gate in front of every write to a system of record; a rubric or eval that scores the agent's work before anyone relies on it; and a named person who is accountable when the agent is wrong. The tooling is the easy part — the frameworks, the models, and tool use over the Model Context Protocol are all mature. The discipline is the product. We say this as operators, not observers: agentic systems run our own company. A work-orchestration MCP server coordinates mission boards and daily operations across our portfolio, an agentic HR platform runs onboarding and compliance with human judgment kept for the decisions, and a self-evolving software platform adapts its own code as the business it serves changes. What follows are the lessons that only showed up once those systems were carrying real work — the things you learn by operating agents, not by reading about them.

The agent is the easy part. The boundary is the work.

Standing up an agent that can plan and act takes days now. What takes real engineering is the boundary around it: which systems it may read, which it may write, what it is supposed to do when it is uncertain, and how its work gets checked. Most agent failures are not model failures — they are boundary failures. An agent handed a broad tool because scoping a narrow one was inconvenient. An escalation path that existed in the design document but not in the tools. A write permission granted for a pilot that nobody revoked. When we build an MCP server, every tool it exposes is scoped to the narrowest action that still does the job, because every tool an agent holds widens what a bad plan can do. The model's judgment is probabilistic; the tool surface is the part you actually control, so that is where the control has to live.

Approval-before-write is a design principle, not a training wheel

The most load-bearing pattern in our systems is that reads are generous and writes are gated. An agent can research, summarize, draft, and propose freely; the moment it wants to change a system of record — a schedule, a compliance status, a line of code — a human approves the change before it lands. Teams usually treat that gate as scaffolding to remove once the agent proves itself. Operating agents taught us the opposite: the gate is what makes the autonomy adoptable at all. It produces an audit trail of human-confirmed changes, it keeps accountability with a person instead of a probability, and it costs one click. Our agentic HR platform runs its routine motion autonomously precisely because the decisions stay human. Autonomy in execution is safe when acceptance is not autonomous.

Autonomy is for the work, not for the acceptance of the work. An agent can plan, execute, and propose; a person decides what counts as done.

Without rubrics, agents drift into confident mediocrity

An agent with no scored definition of good does not fail loudly; it degrades quietly toward plausible. The output stays fluent, the formatting stays clean, and the substance erodes — and because each individual piece looks reasonable, nobody notices until the aggregate is wrong. The fix is the same eval discipline we apply to any model system: a versioned set of real cases with a defined correct outcome, a rubric that turns 'seems fine' into a score, and a rule that prompt edits, model swaps, and tool changes all re-run the evals before they ship. Monitoring tells you the agent is up. Evals tell you it is still good. They are different instruments, and production agents need both.

  • An agent that is right most of the time teaches people to stop checking it. Review has to be structural — built into the workflow — because human vigilance decays at exactly the rate the agent earns trust.
  • Ambiguity does not announce itself. An agent facing an unclear situation acts on its best guess unless an explicit escalate-to-human path exists in its tools, not just in the architecture diagram.
  • Behavior changes without errors. A model version bump or an edited prompt shifts outputs with nothing in the logs — only an eval run catches it before the people downstream do.
  • Capability grows appetite. As agents get better, they attempt more, and a per-task budget for cost and scope is what keeps an ambitious plan from becoming an expensive one.

Why we do not vibe-code — and why agents raise the stakes

Our developers do not vibe-code applications. Before AI writes a line, requirements become structured specs — data models, acceptance criteria, and the success metrics the delivery will be measured against. AI provides velocity, the spec provides direction, and a senior engineer reviews every change before it merges; an AI security pass and a human verification pass run before anything ships. Agents make that discipline more important, not less. Our self-evolving software platform adapts its own code as the business it serves changes, and it stays trustworthy for exactly one reason: its changes go through the same gates as human-written code. Code review does not care who the author was. The day an agent's output skips the review a person's output would face, you no longer have an engineering process — you have an unmonitored dependency with commit access.

The methodology these lessons produced

Our delivery methodology — Survey, Blueprint, Build, Commission — is these lessons in sequence. Survey draws the data boundary before anything else. Blueprint writes the rubric, the eval set, and the cost budget before a model is chosen. Build runs the spec-driven, human-gated SDLC described above. Commission ships with observability and keeps the approval gates in place instead of quietly removing them at launch. We recommend it to clients because it is what we run on ourselves.

If you are planning your first production agent, the useful first conversation is not about which model or framework to pick. It is about where the approval gates sit, who owns the rubric, and which writes a human must confirm. We have that conversation from the operator's side of the table — the systems described here run our own operations every day.

Have a project like this?Book a working session
Newsletter

Field notes from the bench

Occasional writing on how we build — analytics, launch discipline, AI engineering. No fluff.

No tracking · Unsubscribe anytime · ~Monthly