Demos love tools: the model calls search, updates a row, posts to Slack, audience applauds. Production asks who authorized that row update, what gets logged, and what happens when the tool returns 500.
Anthropic’s Model Context Protocol (MCP) standardises how clients discover and call tools. Claude API tool use is the runtime contract. At PAI Technologies we use both: MCP for integrator ergonomics, explicit FastAPI/LangGraph boundaries for what users touch.
This article is how we go from “cool MCP server” to audited agent behavior on PlaywithDB-style NL2SQL and client automation work.
MCP vs in-app tools: when each fits
MCP shines for developer velocity: one server exposes resources and tools to Claude Desktop, Cowork, or internal harnesses without bespoke glue per client.
User-facing products usually wrap tools behind your API: you validate session, tenant, and rate limits before any tool runs, MCP endpoints are not dropped straight into the browser.
LangGraph nodes call typed Python functions with shared auth context; MCP can mirror those same operations for internal debugging: single implementation, two surfaces.
Design tools narrow and explicit

Prefer `query_playground_readonly` over `run_sql` with no guardrails. PlaywithDB separates read paths from destructive ops; production agents inherit that split.
JSON schemas with enums, max lengths, and required fields reduce model creativity in dangerous places. Validate server-side even if the model “followed” the schema.
Return structured errors the model can retry: “tenant not found” not stack traces. Log full detail server-side only.
Auth, tenancy, and secrets
Tools never receive long-lived API keys from the model. The server holds credentials; the tool call carries user/session scope only.
EU deployments: subprocessors documented; tool calls that leave the region (external search, email) gated by policy flags per tenant.
Rotate keys without retraining, tool implementations read from secret managers; agent prompts stay stable.
Observability and evals

Log: model version, tool name, arguments hash (not always raw PII), latency, success, user id, conversation id. Gordon-scale clients ask for retention, answer in architecture docs.
Golden sets include tool-selection cases: when the model should refuse, call tool A vs B, or escalate. CI fails deploy when tool-routing accuracy drops.
Replay production failures into eval sets weekly: the best tool prompts come from real misroutes, not imagination.
Reliability: retries, idempotency, timeouts
Idempotency keys on write tools, retries are inevitable. n8n-style double-send is worse when the tool charges money or emails customers.
Timeouts per tool class: search can be slow; payment cannot hang forever. Circuit-break hot dependencies.
Human approval steps for irreversible actions: refunds, mass email, firmware flash, LangGraph interrupt patterns work well here.
Shipping checklist
- Tool list reviewed with security; no admin tools on user-facing agents.
- Staging environment with fake tenants and recorded eval runs.
- Runbook for disabling a single tool without taking down chat.
- Cost caps per session (model + tool API charges).
- Handover: MCP server repo or FastAPI module, schemas, and eval scripts.
