Demo agents feel magical: five tools, one prompt, a curated question. Production agents fail on auth, ambiguous tool schemas, runaway loops, and legal review of what gets logged. We have shipped Claude-based agents across RAG, ops automation, and EU-hosted products. This is the pattern we repeat.
PAI Technologies builds agents with Claude API tool use, LangGraph when state and human approval matter, FastAPI for boundaries, and eval gates before side effects. Gordon, PlaywithDB-adjacent workflows, and Upwork client builds share the same discipline: measure before you scale traffic.
This article covers architecture, tools, evals, and handover, not prompt tricks alone.
Agent architecture that survives week three
Separate orchestration (LangGraph or explicit state machine) from tool implementations (HTTP, SQL, vector search). Tools return structured JSON; the model reasons over summaries, not raw 10MB payloads.
Every side effect (email, ticket update, payment) sits behind an application service with idempotency keys: the LLM proposes, code commits.
Human-in-the-loop checkpoints for high-risk actions: refunds, external emails, production config changes. EU clients often require this by policy, not preference.
Claude tool use: design rules

Prefer many small tools with clear names and Zod/Pydantic schemas over one “do_everything” function. The model routes better when tools mirror how engineers think about capabilities.
Document failure modes in tool descriptions: “returns 404 if order not found” helps refusal and retry behaviour.
Cap iterations and spend per session. LangGraph interrupts when limits hit; log why for tuning.
Version tool schemas with the agent release, breaking changes without version bumps cause silent regressions.
LangGraph vs simpler loops
Use LangGraph when you need branching, persistence, human approval, or replay for debugging. Simple retrieve-then-answer RAG may need only two nodes.
Checkpointing helps EU audit scenarios: reconstruct what the agent knew before an action. Store checkpoints with retention aligned to GDPR policy.
PlaywithDB-style agents add governed SQL tools: schema snippets in context, row-level policies enforced in code, not in prompts.
Evaluation before production traffic

Golden conversations: user messages, expected tool calls (or forbidden tools), and properties of final answers. Run nightly in CI.
Adversarial tests: prompt injection via retrieved docs, requests for secrets, cross-tenant data probes.
Track tool error rate, average steps per task, cost per successful outcome, not only user thumbs.
Block deploy when metrics regress; keep human review samples for long-tail drift.
Observability and compliance
Structured logs: run ID, model version, tools invoked, retrieval chunk IDs, latency, token cost. Redact PII at log boundary.
For Gordon-class EU hosting, run agents in-region with subprocessors listed for client legal review. Retention jobs delete old runs per policy.
Expose a support view: replay with redaction for debugging without giving every engineer full prompt access.
Real client patterns we see
Support copilot: RAG + ticket update tool with human approve on send, ships in one three-to-four-week cycle with eval on 100 historical tickets.
Internal ops agent: hybrid n8n trigger → FastAPI LangGraph core → CRM, see our automation article for the split.
Data analyst agent: PlaywithDB or custom NL2SQL with read-only roles and citation of tables used, compliance loves attributable answers.
We hand over repo, eval scripts, runbooks, and Loom walkthrough. You own the agent, PAI’s default on every engagement.
Send a brief with tools you need, data sensitivity, and region; we will propose a vertical slice and honest timeline.
