Agent Infrastructure

AgentCore V2's Two-Second Result Is a Cold-Start Test, Not a Finished Agent Task

By Kaleido Field Staff ยท September 28, 2026

Keep startup latency separate from task latency

The agent in AWS's cold-start test does not call a model or a tool. The September 18 AgentCore runtime launch reports roughly two-second P75 starts across five container sizes using an echo workload. That is a platform-start measurement, not the time to complete a real agent task.

Citation-ready: AWS's AgentCore V2 cold-start result measures an empty echo agent over a specified cross-Region client path; it does not establish two-second completion of model or tool work.

Evidence boundary: AWS-reported launch and company-run test, not an independently replicated benchmark or a production task SLA. Future pricing and compute options remain future capabilities. Publication note: Prepared for September 21, delayed by a deployment failure, and published September 28 after source revalidation. Original event dates are retained; this is not a new September 28 announcement.

AWS chart comparing original and new runtime P75 cold-start latency across 200 MB to 2 GB container images
Image source: AWS; company-produced echo-agent cold-start chart, including client path and sample count. Used for editorial coverage of runtime latency and resource accounting desk.

What happened and why it matters

The launch improves a measurable part of the wait, but model inference, tools and the agent loop remain outside its test workload.

Primary evidence

Primary reference: AWS AgentCore runtime launch and test-method article. Kaleido Field checked the event date and the article's attributed facts against this source.

Source check
Source dateSeptember 18, 2026
Checked by Kaleido FieldSeptember 28, 2026, CST
Source functionagent infrastructure -> startup measurement and workload-dependent economics

The chart has useful controls and important exclusions

AWS compares five image sizes from 200 MB to 2 GB. The plotted V2 P75 values range from 1.9 to 2.0 seconds, including the client's cross-Region round trip. The source says the agent's own echo code contributes little of that time.

For a production trial, record startup separately from inference, tool calls and retries. A faster start can matter to an interactive user without being the largest portion of the total wait. Use the same request mix when comparing versions.

Elastic memory changes the accounting question

The new runtime uses snapshots and reclaims released or cold memory instead of holding the session's peak. AWS says the rate is higher but the memory footprint is often lower; that is not a guarantee of a lower bill for every workload.

Compare actual resource use and charges over a representative run. Baseline discounts, larger resources, x86 support and additional lifecycle controls are listed as coming next, not as part of the verified launch. Our inference-measurement report explains why the workload definition belongs beside every performance result.

Evidence boundary

AWS-reported launch and company-run test, not an independently replicated benchmark or a production task SLA. Future pricing and compute options remain future capabilities.

Reader briefing

Keep the source trail in view.

One concise email when a model, benchmark, or visual-intelligence claim materially changes.

FAQ

Does the AWS echo test include the time spent asking a model or using a tool?

No. The stated echo agent returns its input without calling models or tools.