How Postman Narrows 170 Agent Tools to About 15
The Agent Mode design uses dynamic tool selection, schema-aware reads, purpose-built context, approval gates and tiered prompt caching instead of exposing every capability at once.
Edited by Tyronne Panaino
Postman and AWS published an engineering account on October 9 explaining how Postman designed Agent Mode for a mature API platform. The disclosed architecture combines dynamic tool selection, schema-aware reads, purpose-built context handlers, user approval before state-changing actions, and Amazon Bedrock routing and prompt caching.
The account matters to teams trying to move agents beyond demonstrations and into products with large tool surfaces and years of interface assumptions. It is also first-party evidence: the joint AWS and Postman case study describes Postman's own implementation, but provides no public repository or independent reproduction of its operational results.
A large tool catalog became a context problem
Postman says its early design favored small, precise tools for individual actions. That improved control at first, but longer workflows required many round trips through the model. The company reports that tool-selection errors increased when the visible catalog grew beyond roughly 40 tools, including calls to nonexistent tools and incorrect arguments.
Its current pattern is selective exposure. A root agent searches embeddings for a catalog of more than 170 tools, narrows the visible set to about 15 that are relevant to the request, and hands them to a context-isolated sub-agent. The practical lesson is that adding capabilities and showing all of them to the model are separate decisions. A smaller task-specific menu can reduce competition between superficially similar actions.
Those figures are Postman's measurements, not a universal threshold. A different product may have clearer schemas, different model behavior or tasks that need a wider set. Teams should measure selection errors and task completion in their own catalog rather than copying the 40-tool observation as a fixed limit.
Schema-aware reads replace many narrow tools
For structured product data, Postman says it consolidated several narrow views into a single query tool backed by database schemas. The agent can compose queries over fields such as service uptime, test results and endpoint latency instead of requiring a separate tool for every question.
That shifts engineering work from multiplying tool definitions to designing a readable data model and enforcing access boundaries around the query surface. It can broaden the questions an agent can answer, while also creating a need to validate generated queries, limit costly operations and prevent the read path from exposing data outside the user's permissions. The source demonstrates the architecture choice, not an independent security assessment of every query it can produce.
Postman also reports that missing or incomplete context caused more failures than missing capabilities. The team built dedicated handlers that reduce each product entity to the information the agent needs, rather than serializing objects designed for the user interface. Open-ended user content can still crowd the context window, so truncation and expansion remain active design problems.
Approval and guardrails sit outside model choice
Agent Mode requires user approval before actions that change application state, according to the case study. Postman also says tools are scoped to the task and that enterprise administrators can enable Amazon Bedrock Guardrails to redact personally identifiable information before a request reaches the underlying model.
These are different controls. Dynamic tool selection decides which actions the model can see for a task. Approval decides whether a proposed state change proceeds. Redaction addresses selected data before inference. None of them alone proves that a tool is safe, that an approval is informed, or that every sensitive field will be detected. Production teams still need authorization checks, logs, monitoring and tests for rejected or malformed actions.
Regional routing and two cache lifetimes support the runtime
Postman says Agent Mode uses Bedrock cross-Region inference profiles. A geographic profile can keep processing within its supported geography, while a global profile can reach a wider set of regions for throughput. The case study is explicit that Bedrock processes the requests in eligible AWS regions rather than inside Postman's own environment.
The company also reports configuring zero data retention for supported models, while warning that availability and behavior depend on the selected model. That qualification makes retention a model-by-model verification task, not a blanket property of every Agent Mode request.
For repeated context, the design uses a one-hour cache checkpoint for the relatively stable core, including system behavior and core tool definitions, and a five-minute checkpoint for more variable conversation and selected context. The longer-lived checkpoint appears first. Postman presents this as a way to avoid repeatedly processing the same prefix; the source does not publish independent latency or cost results for the complete workload.
What the evidence establishes
The case study provides a concrete design pattern for reducing tool sprawl, separating product data from interface state, managing context and placing approval, regional routing and caching around a production agent. It does not disclose Postman's proprietary implementation, task-level success rates, false-action rates, traffic volumes, cost savings or an external audit of data handling. The reference to 40 million developers describes Postman's community scale and should not be read as an Agent Mode usage count.
The next useful checkpoints are reproducible evaluations across larger tool catalogs, quantified approval and tool-selection failures, independently verified retention behavior, and evidence that the approach remains reliable as the product and its knowledge base change.
Status
Learning. AWS and Postman document the architecture and reported measurements. Internal confidence is medium because the evidence is a joint first-party account and the production implementation is not available for independent inspection.
Sources
Update note: Last reviewed 2026-10-11. We will revise this post if Postman publishes reproducible evaluations, implementation artifacts or independent operational evidence.
Sources
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.