All Case Studies
AI Architecture
Reference Architecture

Production AI Architecture at Scale

Reference architecture and operating model for production AI across a multi-cloud enterprise: shared entry points, cost attribution, onboarding standards, and EU AI Act evidence built into the delivery path.

AWS Summit Madrid 2025: Scalable AI Adoption at Iberdrola (speaker, ES)
Production AI Architecture at Scale

Executive Outcome

01

New teams onboarded through the standard path instead of building bespoke platform components from scratch, which cut both duplication and time to first deployment.

02

Model usage and cost attribution became visible across business units for the first time, which changed how leadership made capacity and investment calls.

03

Governed entry points became the default route for new GenAI and agentic workloads — shadow AI dropped, and security enforcement got consistent without turning into a bottleneck.

Engagement focus

Reference architecture and operating model for production AI in a federated multi-cloud enterprise: shared controls with distributed ownership.

What this covers
  • Shared platform architecture with explicit plane boundaries and entry points
  • Cost attribution, observability, and onboarding standards across business units
  • Governance and compliance evidence embedded into delivery gates with progressive adoption

Context

A European energy group had multiple teams running independent AI experiments across several cloud accounts and providers. Each team was solving the same platform problems — gateways, identity, logging, cost tracking — on its own, in parallel, mostly without knowing the others were doing it too. The result was duplicated infrastructure, zero central visibility into model consumption, and no consistent way to enforce security or governance at scale. Centralizing everything wasn't the goal; I needed a shared architecture teams could adopt at their own pace, with clear ownership boundaries and enough room for local adaptation that it wouldn't just get bypassed.

The Challenge

  • 01Teams were reinventing gateways, identity, and observability independently across business units — duplicated effort, fragmented standards.
  • 02There was no cost visibility or consumption attribution. Model usage was invisible to finance and to the platform team alike.
  • 03Delivery standards were inconsistent enough that enforcing security baselines or governance meant either blocking teams or looking the other way.
  • 04Shadow AI risk kept growing as teams found unmanaged paths to models with no central telemetry watching.

Approach

  • →Drew explicit plane boundaries in the reference architecture — teams own their workloads, the platform owns shared controls, and the line between the two is documented, not implied.
  • →Stood up standard entry points for model and tool interactions with shared routing, telemetry, and cost attribution.
  • →Wrote onboarding standards and an ownership model so new teams could adopt the platform through a repeatable path instead of a bespoke negotiation each time.
  • →Embedded governance and compliance evidence into the release gate design itself, so it's part of how something ships rather than a separate review that happens afterward.
  • →Designed for progressive adoption — teams onboard on their own timeline, with sensible defaults and exception paths instead of a mandate handed down.

Key Considerations

  • Standardized access paths trade some team autonomy for consistent governance and lower overall overhead — worth it, but it felt like friction early on.
  • Centralizing model access creates a platform dependency, which means clear service expectations and an actual incident response plan, not an afterthought.
  • Mandatory onboarding adds early friction, but the alternative is teams quietly rebuilding platform controls on their own.
  • Progressive adoption means some teams keep running on legacy patterns during the transition — that needs an explicit migration timeline or it never ends.

Alternatives Considered

  • ✕A library-only approach doesn't give you central enforcement or any visibility into model consumption — teams can simply not use it.
  • ✕Betting on a single vendor creates lock-in and limits control over identity and access patterns down the line.
  • ✕A central mandate without adoption support generates resistance and shadow workarounds fast in a federated organization — I've seen it happen more than once.
Representative Artifacts
01Reference Architecture (plane boundaries, entry points, shared services)
02Platform Map (identity, routing, telemetry, cost attribution)
03Onboarding Pack (templates, checklists, ownership expectations)
04Ownership Model (platform vs. application team responsibilities)
05Release Gate Definitions (governance and compliance evidence embedded)
06Migration and Adoption Roadmap (progressive onboarding, exception paths)
Acceptance Criteria

New teams onboard through the standard path without bespoke platform intervention.

Telemetry captures interaction traces consistently for security and cost attribution across business units.

Ownership boundaries reflected in delivery standards and review checkpoints.

EU AI Act evidence generated as part of the standard release gate process.

Exception paths documented and governed for teams with legitimate local adaptation needs.

Continue Exploring

Other Case Studies