Migration Readiness for AI Platforms: What to Prove Before You Trust the TCO

Written by

in

Teams moving an AI platform to the cloud often go straight from “here is what we run today” to “here is the AWS architecture and its three-year cost.” That shortcut is tempting, because the current system already works and the flows are well understood. It is also how a development-stage topology gets translated, service for service, into a production design that inherits every one of its weaknesses.

This article sets out what a Migration Readiness Assessment (MRA) should establish for a multi-use-case AI platform, what the target architecture must then do with that evidence, and how the three-year Total Cost of Ownership (TCO) should be treated until the two are reconciled.

Key principle: current-state flows should inform requirements. The MRA should expose readiness gaps and production needs. The target architecture should then make explicit design decisions that close those gaps. The current implementation is evidence of behavior and dependencies, not a design pattern to copy.

1. What an MRA Is, and What It Is Not

AWS describes a Migration Readiness Assessment as a process that shows how far along an organization is in its cloud journey, what its current strengths and weaknesses are, and what action plan will close the identified gaps. It is based on the AWS Cloud Adoption Framework (CAF) and its six perspectives: business, people, governance, platform, security, and operations.

For an AI platform, that means the technical assessment has to do more than document the happy path of each use case. Three artifacts should stay distinct:

  • The current state and use-case flows: what runs today and what each use case does.
  • The MRA: what is missing between today’s state and production readiness.
  • The target architecture: the design decisions that close those gaps.

2. What the Assessment Must Establish

  • Validate the current state. Confirm what is implemented, what is prototype or development-stage, what is feature-gated, and what is not operational at all.
  • Capture functional requirements. Use the use-case flows to establish what users, agents, applications, and supporting services must do. Do not assume today’s implementation mechanism is tomorrow’s AWS mechanism.
  • Identify dependencies. Record use-case-to-use-case and platform dependencies, such as document ingestion and indexing feeding search and research workloads.
  • Identify shared capabilities. Treat common capabilities (search, knowledge stores, agent runtime, security and governance, review tooling) as platform concerns instead of rebuilding them for every use case.
  • Turn limitations into production requirements. Translate each gap into explicit non-functional requirements (NFRs): availability, scale, resilience, disaster recovery, security, compliance, observability, performance, operability, environment separation, deployment automation, and cost governance.
  • Separate facts from assumptions. Record unknown workload volumes, concurrency, growth, retention, recovery targets, and service-level expectations as assumptions or discovery actions. Do not embed them silently in the architecture or the cost model.
  • Create a gap-closure plan. For each material gap, state the remediation, the owner or decision required, and whether it affects the target architecture, the migration sequence, or the cost.
  • Feed the architecture and the TCO. The output should be a set of traceable requirements and constraints from which both the design and the cost model can be derived.

One rule deserves emphasis: where nobody has defined a recovery time objective, target concurrency, latency goal, or SLA, obtain or agree those values during discovery. Do not invent them inside the target architecture.

3. Readiness Gaps That Should Shape the Target State

Prototype-stage AI platforms tend to share a recognizable set of gaps. Each one should map to a concrete requirement rather than a generic aspiration.

AreaTypical gapTarget-state requirement
Resilience and availabilityNo production-grade redundancy or failover has been demonstrated.Define the availability model, failure domains, Multi-AZ needs, health and recovery behavior, and service dependencies.
Scale and performanceCapacity has only been observed at small scale, well below expected production load.Define user, concurrency, ingestion, corpus, token, and tool-call assumptions, plus a load-testing approach and scaling model.
Disaster recoveryBackups exist, but recoverability has not been demonstrated.Define recovery time and recovery point objectives (RTO and RPO), restore and replication design, recovery testing, and a regional strategy where justified.
Software lifecycle and environmentsLimited separation between environments and informal release practices.Define dev, test, and prod separation, CI/CD, infrastructure as code, release controls, and rollback.
Security and governanceProduction controls not yet designed explicitly.Define identity, authorization, tenant and workspace isolation, secrets, network controls, audit, logging, encryption, data residency, and guardrails.
Observability and operationsNo agent-level production tracing or formal operational practice.Define agent traces, tool and model telemetry, dashboards, alerts, incident response, service-level objective (SLO) inputs, and operational ownership.
Agent architectureDifferent use cases orchestrate agents in different ways.Decide what is shared and what is use-case-specific, and define a common runtime pattern without needlessly locking in one framework.
Shared platformSeveral use cases depend on common search, agent, knowledge, governance, and review capabilities.Design shared services and their boundaries before treating each use case as its own infrastructure stack.

4. What the Target Architecture Must Do

The target architecture is a production design derived from validated requirements and gaps. It is not a one-for-one cloud translation of the development implementation. Specifically, it should:

  • State its principles and boundaries. Show the application and API layer, agent runtime and orchestration, tools, retrieval, memory and state, data, identity and security, observability and evaluation, and document processing. Explain what remains your own intellectual property and what becomes a managed cloud capability.
  • Define the shared platform first. Establish the common services and controls, then overlay each use case’s specific flow. This avoids ending up with one duplicated architecture per use case.
  • Record an explicit decision for each managed service. For a managed agent platform such as Amazon Bedrock AgentCore, that means Runtime, Gateway, Identity, Memory, Observability, Evaluations, Code Interpreter, and any others that are relevant. Mark each one Use, Do not use, or Later, with the rationale. These services are composable, so none should be assumed mandatory. For one worked example of a target design built around AgentCore Runtime, see AI Startup Architecture on AWS AgentCore: A Reference Blueprint.
  • Stay framework-neutral where it makes sense. Name a web framework or agent orchestration library only where it is genuinely a target-state constraint. Otherwise draw logical service boundaries.
  • Decide how agent workflows relate to each other. Whether two workflows stay separate, share a runtime and tool layer, reuse retrieval capabilities, or converge in specific areas should be an intentional, documented decision.
  • Enforce access control deterministically. User, organization, and workspace boundaries, tool authorization, least privilege, data protection, and auditing must be enforced independently of model reasoning. Never rely on a model to keep tenants apart.
  • Show mechanisms and acceptance criteria for every NFR. Statements like “scalable” or “well architected” mean nothing without them.
  • Include the operational architecture. Environment strategy, infrastructure as code, CI/CD, telemetry, agent tracing, evaluation, incident response, backup and restore, and cost controls all belong in the design, not in a post-launch to-do list.
  • Be traceable. Every major component should link back to a functional requirement, shared capability, NFR, gap remediation, or explicit architecture decision.

A Decision Record for Every Material Component

Where a current component overlaps with a managed cloud capability, capture the choice in a short decision record. Each record should contain:

  • The requirement or capability being satisfied
  • The current implementation and its dependencies
  • The known gap or production NFR
  • The target service or pattern, and the alternatives considered
  • The reason for selection or retention
  • Whether the scope is shared or use-case-specific
  • Security and data-residency implications
  • Availability and scaling assumptions
  • The cost and TCO impact
  • Open questions that need customer validation

5. How the 3-Year TCO Fits In

An existing three-year TCO should be treated as a baseline until it is reconciled against the agreed production target architecture. That does not mean rebuilding the cost model from scratch. Instead, classify each existing cost line as Retain, Resize, Replace, Remove, or Add based on the target-state decisions. The chain of evidence should read left to right, with each step traceable to the one before it.

From current state to a validated 3-year TCO Five connected steps: current state and use-case flows, requirements and dependencies, gaps and production non-functional requirements, target state on AWS and AgentCore, and a validated three-year TCO. Each TCO cost line is classified as retain, resize, replace, remove, or add. Current state & use-case flows Requirements & dependencies Gaps & production NFRs Target state (AWS / AgentCore) Validated 3-year TCO Every cost line is classified: Retain · Resize · Replace · Remove · Add
The chain of evidence from current state to a validated three-year TCO.

6. Minimum Acceptance Criteria

Before calling a target architecture complete, check that:

  • Every use case is represented at requirements level, even where detailed designs are phased.
  • Shared platform capabilities and use-case dependencies are explicitly mapped.
  • Every material gap has a target-state remediation or a documented exception.
  • Production NFRs have measurable targets, or are recorded as unresolved discovery items.
  • Service choices are justified on their merits, not selected only because they mirror current components.
  • The orchestration relationship between agent workflows and the shared platform is explicitly decided.
  • Security, governance, observability, resilience, disaster recovery, and operations are architectural concerns, not post-design prerequisites.
  • The three-year TCO is reconciled to the target architecture and names its assumptions and sensitivity drivers.
  • Open decisions are visible and do not masquerade as confirmed facts.

7. A Recommended Working Sequence

  1. Validate the current-state facts and unknowns with the platform team.
  2. Confirm functional requirements for every use case.
  3. Map dependencies and shared platform capabilities.
  4. Convert gaps into production NFRs and gap-closure actions.
  5. Agree target-state principles and the shared platform.
  6. Design use-case overlays, starting with the priority use cases.
  7. Record managed-service decisions and their rationale.
  8. Reconcile the existing TCO using Retain, Resize, Replace, Remove, and Add.
  9. Validate the architecture, NFR coverage, and TCO assumptions with the customer or business owner.

Conclusion

A target architecture and a three-year TCO are only as trustworthy as the readiness evidence behind them. Treat the current implementation as evidence, not as a blueprint. Convert gaps into measurable requirements, design shared capabilities before per-use-case detail, and keep every open assumption visible until someone has actually validated it. To go further, see AI Startup Architecture on AWS AgentCore: A Reference Blueprint for a worked target design, and Cloud Migration: AI Startups vs. Legacy Enterprises for how migration priorities differ by workload type.

Sources