Back to blogAI & Software

Building Production AI Agents with Claude: API, Tools, and Evals

Tool use, LangGraph, evaluation gates, and handover patterns we use for client agents on Claude, from pilot to production in focused cycles.

Share this article

AI and code workspace for production Claude agents: PAI Technologies blog

Demo agents feel magical: five tools, one prompt, a curated question. Production agents fail on auth, ambiguous tool schemas, runaway loops, and legal review of what gets logged. We have shipped Claude-based agents across RAG, ops automation, and EU-hosted products. This is the pattern we repeat.

PAI Technologies builds agents with Claude API tool use, LangGraph when state and human approval matter, FastAPI for boundaries, and eval gates before side effects. Gordon, PlaywithDB-adjacent workflows, and Upwork client builds share the same discipline: measure before you scale traffic.

This article covers architecture, tools, evals, and handover, not prompt tricks alone.

Agent architecture that survives week three

Separate orchestration (LangGraph or explicit state machine) from tool implementations (HTTP, SQL, vector search). Tools return structured JSON; the model reasons over summaries, not raw 10MB payloads.

Every side effect (email, ticket update, payment) sits behind an application service with idempotency keys: the LLM proposes, code commits.

Human-in-the-loop checkpoints for high-risk actions: refunds, external emails, production config changes. EU clients often require this by policy, not preference.

Claude tool use: design rules

Developer writing code for Claude API tool integrations
Small, composable tools with strict schemas outperform mega-tools.

Prefer many small tools with clear names and Zod/Pydantic schemas over one “do_everything” function. The model routes better when tools mirror how engineers think about capabilities.

Document failure modes in tool descriptions: “returns 404 if order not found” helps refusal and retry behaviour.

Cap iterations and spend per session. LangGraph interrupts when limits hit; log why for tuning.

Version tool schemas with the agent release, breaking changes without version bumps cause silent regressions.

LangGraph vs simpler loops

Use LangGraph when you need branching, persistence, human approval, or replay for debugging. Simple retrieve-then-answer RAG may need only two nodes.

Checkpointing helps EU audit scenarios: reconstruct what the agent knew before an action. Store checkpoints with retention aligned to GDPR policy.

PlaywithDB-style agents add governed SQL tools: schema snippets in context, row-level policies enforced in code, not in prompts.

Evaluation before production traffic

Analytics dashboard for AI agent evaluation metrics
Golden sets, tool-call assertions, and refusal tests block bad releases.

Golden conversations: user messages, expected tool calls (or forbidden tools), and properties of final answers. Run nightly in CI.

Adversarial tests: prompt injection via retrieved docs, requests for secrets, cross-tenant data probes.

Track tool error rate, average steps per task, cost per successful outcome, not only user thumbs.

Block deploy when metrics regress; keep human review samples for long-tail drift.

Observability and compliance

Structured logs: run ID, model version, tools invoked, retrieval chunk IDs, latency, token cost. Redact PII at log boundary.

For Gordon-class EU hosting, run agents in-region with subprocessors listed for client legal review. Retention jobs delete old runs per policy.

Expose a support view: replay with redaction for debugging without giving every engineer full prompt access.

Real client patterns we see

Support copilot: RAG + ticket update tool with human approve on send, ships in one three-to-four-week cycle with eval on 100 historical tickets.

Internal ops agent: hybrid n8n trigger → FastAPI LangGraph core → CRM, see our automation article for the split.

Data analyst agent: PlaywithDB or custom NL2SQL with read-only roles and citation of tables used, compliance loves attributable answers.

We hand over repo, eval scripts, runbooks, and Loom walkthrough. You own the agent, PAI’s default on every engagement.

Send a brief with tools you need, data sensitivity, and region; we will propose a vertical slice and honest timeline.

Share this article