Stop Dreaming, Start Engineering: Cloud Patterns for Production AI

A presentation at KCDC 2026 in September 2026 in by Jeremy Meiss

Slide 1

Slide 1

Stop Dreaming, Start Engineering Cloud Patterns for Production AI Jeremy Meiss — Tech Solution Architect

Slide 2

Slide 2

“ How many of you have shipped an AI demo, feature, or new application that quietly died?

Slide 3

Slide 3

80%+ of AI projects fail — twice the rate of conventional IT RAND, 2024. 65 practitioner interviews — citing an outside estimate. ☀ 1 / 48

Slide 4

Slide 4

Two more studies 95% of GenAI pilots show no measurable P&L return The GenAI Divide. MIT Project NANDA, 2025. Preliminary, not peer-reviewed. Two more studies 28% of AI use cases in infrastructure & operations meet ROI expectations Gartner, April 2026. 782 I&O leaders. ☀ 1 / 48

Slide 5

Slide 5

Jeremy Meiss Tech Solution Architect · WWT DevOpsDays KC organizer CommunityDays KC organizer CDF Ambassador 2025-26 World Wide Technology ☀ 1 / 48

Slide 6

Slide 6

Slide 7

Slide 7

Where the failure actually is The notebook model.fit() converges Accuracy looks great on the holdout set Demo day: it works The system around it Data pipeline — unowned, undocumented Identity boundary — nobody drew it Observability — bolted on after the incident Cost line — nobody modeled it Where the failure actually is ☀ 1 / 48

Slide 8

Slide 8

The story we tell ourselves vs. the reality The story… model works in the notebook! we’re 90% done!!!!! The reality… When the model works in the notebook, you’re maybe 10% done. The other 90% is the engineering we’ve spent 30 years figuring out for every other kind of software + some nrew rules for the parts that are different. The story we tell ourselves vs. the reality ☀ 1 / 48

Slide 9

Slide 9

“ AI is software with new rules.

Slide 10

Slide 10

The six-phase lifecycle Data Processing Business Goal The six-phase lifecycle Problem Framing Model Development Deployment Monitoring ☀ 1 / 48

Slide 11

Slide 11

Use-case drift Week 1 Month 3 Month 6 “Reduce complaints by 30%” Which vector database? Nobody mentions complaints anymore One of RAND’s five root causes: stakeholders miscommunicating the problem Use-case drift ☀ 1 / 48

Slide 12

Slide 12

Well-Architected, for AI Operational Excellence Security Reliability MLOps automation Model, data, prompt Recovery from drift Performance Cost Sustainability Purpose-built silicon Inference is your bill Measurable now AWS ML Lens · GCP AI/ML Framework · Azure Well-Architected AI workload guidance Well-Architected, for AI ☀ 1 / 48

Slide 13

Slide 13

Three competitors. Same six answers. AWS GCP Azure All three arrived at the same six Well-Architected pillars — independently. AWS ML Lens · GCP AI/ML Framework · Azure Well-Architected AI workload guidance Convergence ☀ 1 / 48

Slide 14

Slide 14

“ Regulation is an architectural input now. Not a Phase-6 checkbox.

Slide 15

Slide 15

The EU AI Act. Two dates, not one. August 2, 2026 Article 50 transparency obligations December 2, 2027 High-risk regime GPAI penalty powers Human-in-the-loop checkpoints Fines up to €15M or 3% of global revenue Risk classification, conformity assessment Enforceable 5 weeks ago Delayed 16 months — Digital Omnibus, May 2026 Source: EU AI Act (Regulation (EU) 2024/1689); Digital Omnibus package, May 2026 The clock ☀ 1 / 48

Slide 16

Slide 16

Standards, not just regulations EU AI Act NIST AI RMF ISO/IEC 42001 Buy, build, or hybrid. Same controls. Same paper trail. Standards, not just regulations ☀ 1 / 48

Slide 17

Slide 17

The core problem A unit test Same input The same test, against an LLM → Same input → one deterministic output three different outputs, three runs Pass or fail “Correct” isn’t pass or fail The core problem ☀ 1 / 48

Slide 18

Slide 18

Four patterns for safe rollouts 01 02 03 04 BlueGreen Canary Shadow Gradual, real-user exposure Mirror traffic, discard results A/B Testing Fast rollback, trust your validation Four patterns for safe rollouts Prove business impact ☀ 1 / 48

Slide 19

Slide 19

Shadow deployment, in detail Production model Response returned Shadow model Logged & compared never returned User request mirror Build 1 — production path only Build 2 — shadow duplication added Build 3 — comparison and diff, labeled Shadow deployment, in detail ☀ 1 / 48

Slide 20

Slide 20

The Azure template 01 02 03 04 05 Deploy green Smoke test Mirror Progressive shift Retire blue 0% traffic Invoke green by name The Azure template Live traffic duplicated → 10% 25% 100% → → 50% ☀ 1 / 48

Slide 21

Slide 21

Trace-based eval Production traces Trace store Autorater Regression test suite AgentCore Evaluations (GA June 17, 2026, AWS Summit New York) · Foundry tracebased evaluation · Agent Evaluation on Agent Platform Trace-based eval ☀ 1 / 48

Slide 22

Slide 22

The sequence 01 02 03 Shadow Canary Full cutover Tells you whether the model is broken Tells you whether users tolerate it The sequence ☀ 1 / 48

Slide 23

Slide 23

“ If shadow isn’t in your pipeline, that’s your highest-leverage fix. World Wide Technology ☀ 1 / 48

Slide 24

Slide 24

Why RAG exists Foundation models hallucinate Foundation models are frozen at training cutoff Foundation models don’t know your company RAG solves all three. Why RAG exists ☀ 1 / 48

Slide 25

Slide 25

Four pipelines, one shape Data collection Feature / embedding Vector store Inference orchestration AWS GCP Azure Collection & storage S3 Cloud Storage Blob Storage Embedding & retrieval OpenSearch Vertex AI (Agent Platform) Azure AI Search Four pipelines, one shape ☀ 1 / 48

Slide 26

Slide 26

The managed-knowledge shift 2024 2026 A dev, staring at the four-pipeline diagram The same dev Custom code, everywhere Bedrock Managed KB (AWS Summit NY, Jun 2026) · Foundry IQ (Build 2026) · GCP RAG Engine The managed-knowledge shift One box: managed knowledge base ☀ 1 / 48

Slide 27

Slide 27

Classic vs. agentic retrieval Classic RAG Query: one vector search Agentic retrieval Sources: one Query: LLM decomposes to subqueries Latency: fast Sources: many, parallel Cost: cheap Latency: slow Best for: FAQ, refunds, policy lookup Cost: expensive Best for: complex, multi-source, conversational Classic vs. agentic retrieval ☀ 1 / 48

Slide 28

Slide 28

The default has flipped In 2026, agentic retrieval ships natively on all three clouds. Classic RAG is still a good answer more often than teams think. Per each platform’s current product docs — Bedrock, Foundry IQ, RAG Engine The default has flipped ☀ 1 / 48

Slide 29

Slide 29

“ When the agent retrieves, who is it acting as?

Slide 30

Slide 30

“ Pick your RAG shape deliberately. Managed unless you have a real reason. World Wide Technology ☀ 1 / 48

Slide 31

Slide 31

Wait — Doesn’t MCP Replace RAG? RAG is a technique MCP is a protocol Chunk documents One server per resource Embed them Any compliant client can call it Store in a vector database Works across frameworks and vendors Retrieve relevant chunks at query time Inject into the prompt Says nothing about retrieval logic itself They’re not competing. MCP doesn’t replace RAG — it’s often how you expose a RAG pipeline to an agent. Your vector search doesn’t disappear. It gets wrapped as an MCP Wait — Doesn’t MCPof Replace RAG? into your app. ☀ 1 / 48 server instead hardcoded

Slide 32

Slide 32

What is an agent, actually Model + tool Input → arrow → output Not an agent What is an agent, actually Model in a loop Tools, observations feeding back Agent ☀ 1 / 48

Slide 33

Slide 33

New failure modes Infinite loops Wrong tool called Hallucinated tool Correct tool, wrong argument Correct result, misinterpreted None of these existed in request/response. Your orchestration pattern decides which of these you fight. New failure modes ☀ 1 / 48

Slide 34

Slide 34

MCP: the setup Before Framework A Tool 1 Framework B Tool 2 After Framework C Framework A Tool 3 Framework B Framework C MCP servers Tool 1 Tool 2 Tool 3 Model Context Protocol. Anthropic, November 2024. Open standard. MCP: the setup ☀ 1 / 48

Slide 35

Slide 35

MCP: the numbers 10,000+ 5 Public MCP servers Vendors (currently) shipping native support Anthropic’s own figure, December 2025 Per each vendor’s own product announcements MCP: the numbers Dec 2025 July 2026 Donated to Agentic AI Foundation Largest spec revision since launch Anthropic’s donation announcement, Dec 2025 Per a third-party MCP version tracker, 2026 ☀ 1 / 48

Slide 36

Slide 36

Six orchestration patterns Sequential Routing Parallel Structured pipelines Clear category dispatch Latency reduction, diverse perspectives ReAct Hierarchical (Magentic) Evaluator Tool use, iterative Quality-critical outputs Open-ended tasks Magentic-One: Microsoft Research Six orchestration patterns ☀ 1 / 48

Slide 37

Slide 37

Runtime landscape, current state AWS AgentCore GCP Agent Platform Microsoft Foundry GA Oct 2025. Policy, Evaluations, and managed harness all GA June 17, 2026 (AWS Summit NY). Payments (Coinbase, Stripe) announced. Announced Cloud Next ‘26. Vertex retired. Skill Registry. Memory Bank profiles GA. Agent Framework 1.0 GA. Foundry Agent Service GA. Hosted Agents targeted GA early July 2026. M365 publishing. AWS Summit New York, June 2026 · Google Cloud Next ‘26 · Microsoft Build 2026 Runtime landscape, current state ☀ 1 / 48

Slide 38

Slide 38

Cross-vendor reality Model AWS GCP Azure Claude native, Bedrock available, Model Garden GA June 29, 2026 GPT increasingly available Gemini Open models everywhere primary, expanding primary cross-cloud emerging everywhere everywhere The frontier models have converged closely on most public benchmarks. The platform gap on governance is the more durable one. Claude on Microsoft Foundry: Microsoft & Anthropic, June 29, 2026 Cross-vendor reality ☀ 1 / 48

Slide 39

Slide 39

“ In 2026, agents run everywhere. What varies is who they answer to.

Slide 40

Slide 40

A2A + AP2 Two agents, different organizations, talking directly. A2A — agent-to-agent AP2 — agent payments Watch this space. A2A and AP2: Google, open protocols A2A + AP2 ☀ 1 / 48

Slide 41

Slide 41

The two identity patterns Identity passthrough Runtime isolation Microsoft Foundry + Fabric GCP Agent Identity User’s Entra token flows through the agent SPIFFE cert bound to container, mTLS everywhere Agent inherits the human Agent has its own identity, tied to where it runs The two identity patterns ☀ 1 / 48

Slide 42

Slide 42

OpenTelemetry GenAI conventions gen_ai.system: “anthropic” gen_ai.request.model: “claude-opus-5” gen_ai.usage.input_tokens: 1834 gen_ai.usage.output_tokens: 412 gen_ai.tool.name: “search_knowledge_base” gen_ai.server.time_to_first_token: 0.312 Datadog · Honeycomb · Grafana · LangChain · CrewAI · AutoGen The de facto standard. Not yet formally Stable. As of July 2026 — OpenTelemetry GenAI semantic conventions, Development status

Slide 43

Slide 43

Cost reality Traditional compute Mostly flat once provisioned Inference Linear with usage No cap 78% of companies run 2+ LLM families 3+ jumped from 36% to 59% in one quarter — Databricks, State of AI Agents 2026 Cost reality ☀ 1 / 48

Slide 44

Slide 44

“ Platform choice matters more than model choice.

Slide 45

Slide 45

The nine questions 1. Are your feedback loops wired from monitoring back to framing? 6. Can your agent ever see data its user can’t? Structurally, not policy? 2. Do you have a shadow stage before real users? 7. Are you emitting OpenTelemetry GenAI conventions? 3. Classic or agentic retrieval — deliberate choice? 8. Article 50 enforceable 4. Which orchestration pattern? Can you draw it in 30 seconds? 9. Do you know your per-inference cost and have a routing strategy? 5 weeks — are you ago ready? 5. Are your tools exposed via MCP? The nine questions ☀ 1 / 48

Slide 46

Slide 46

“ The engineering already exists. The standards are here. The regulators are ready. Stop dreaming.

Slide 47

Slide 47

Thank You Jeremy Meiss jerdog.dev /in/jeremy-meiss jerdog jmeiss.me Slides: https://speaking.jmeiss.me