The demo works on a curated dataset. Stakeholders sign off. Then integration, auth, observability, and compliance land on the roadmap, and the “three-month phase two” begins. We built PAI Technologies to avoid that cliff: production path parallel to the pilot, not after it.
Across 30+ engagements and products like PlaywithDB, we run three-to-four-week cycles for a vertical slice that is deployable, measurable, and handover-ready. This playbook is what we use for EU clients (including Gordon for M-agi-c Solutions on eu-central infrastructure), US scale-ups, and India-based product teams.
The timeline is aggressive but repeatable when scope is frozen, evals exist before tuning prompts, and compliance is scoped in week one, not week twelve.
Why most AI pilots stall after the demo
Pilots optimise for the meeting: a slick UI, cherry-picked questions, no auth, synthetic data. Production optimises for failure modes: bad retrieval, PII in logs, rate limits, model outages, and legal review of subprocessors.
Teams that treat the demo as “phase one complete” inherit technical debt in retrieval, evaluation, and deployment. We instead define production acceptance criteria before the first line of feature code.
Gordon, our active EU AI cooking platform engagement for M-agi-c Solutions, Switzerland, follows the same rhythm: staging on region-appropriate infrastructure early, with data flows documented for GDPR review alongside feature work.
Week 1: Freeze scope and pick a vertical slice

Pick one journey that proves value: “analyst asks a question, gets an answer with citations from Postgres,” or “user uploads a recipe, assistant suggests substitutions with sources.” Write a one-page acceptance doc both business and engineering sign.
Choose stack for speed and your constraints: FastAPI or Next.js API routes, Postgres + pgvector or existing search, one model provider to start. PlaywithDB engagements often centre on NL2SQL and governed data access rather than a generic chat shell.
Stand up staging on day two with real anonymised data, not synthetic-only. If anonymisation is hard, that is a week-one finding, not a launch surprise.
- Explicit out-of-scope list posted in the project channel
- Auth model chosen (SSO, API keys, or session) before UI polish
- Golden question set drafted (even 20 rows) before prompt tuning marathons
Week 2: RAG, agents, or data layer with evaluation gates
Whether you use RAG, tool-calling agents, or NL2SQL, you need expected behaviour on a small golden set before you chase prompt elegance. Log every request: prompt, retrieval chunks, latency, token cost.
Add refusal paths for out-of-scope questions, reduces hallucination risk and satisfies legal reviewers faster than disclaimers alone.
Run a lightweight eval nightly; block releases when accuracy or citation rate drops below threshold. For hybrid search (vector + keyword), measure retrieval hit rate separately from answer quality.
Weeks 3–4: Harden, comply, hand over

Production means auth, rate limits, PII handling, structured logging, and a runbook, not just a Docker image. For EU clients we default to region-specific hosting (eu-central-2 Zurich for Gordon-class workloads) and document subprocessors for vendor review.
Healthcare and fintech map controls to HIPAA or PCI-DSS early. Retention policies for prompts and logs are implemented in the same sprint as the feature, not a “compliance phase” months later.
We hand over repos, infra notes, eval scripts, and a short Loom walkthrough. You own the code. That is how three-to-four-week cycles compound into long-term partnerships instead of vendor lock-in.
Stack choices we reach for often
Backend: FastAPI for AI services with clear OpenAPI contracts; Next.js when the team is frontend-heavy and API routes suffice. Agents: LangGraph when state machines and human-in-the-loop matter; simpler tool loops when they do not.
Data: Postgres remains our default for transactional data plus pgvector; hybrid search when exact IDs and SKUs matter. PlaywithDB extends this for teams that need governed natural language over live warehouses.
Observability: request logging with redaction, cost dashboards, and eval hooks. Clients should see p95 latency and cost per query before go-live, not after finance complains.
When three to four weeks is not enough
Fine-tuning, multi-tenant isolation at scale, or novel hardware bring-up may need a second cycle. We say that in the estimate rather than over-promising v1.
If internal IT moves slowly on VPC access or DPO sign-off, calendar time stretches even when engineering is ready, another reason week-one compliance scoping matters.
Send a brief with data, region, and success criteria; we will propose a slice that is honest about what ships in one cycle versus two.
