Back to Community Blog

One connection. 2,000 tools.

authorImage

PayPal Tech Blog Team

Oct 08, 2026

11 min read

featuredImage

How PayPal turned MCP from a common protocol into a governed platform for identity, discovery, certification, and scale.

By the PayPal AITech Cosmos AI Team: Amy Wang, Faiz Mohammed, Suraj Kumar, Olympia Saha, Digvijay Singh, Simon Xu, Divya BH, Srinivas Manoharan

Key figures

  • 6,000 active users
  • 3,500 monthly active users
  • 175 certified servers
  • 2,000+ callable tools
  • ~99% less schema context

Historical point-in-time figures as of May 2026, not public service-level commitments.

An AI assistant can draft an incident update, retrieve a metric, or assemble project context. Inside a large company, every useful answer crosses live systems. The hard part is carrying identity, policy, and auditability with the request.


The platform problem

MCP speaks the language. Enterprises still need the operating model.

Anthropic introduced the Model Context Protocol in November 2024 as an open standard for connecting AI applications to tools and data. MCP later moved under the Linux Foundation's Agentic AI Foundation, giving the protocol a vendor-neutral home.

The standard solves a vital interoperability problem. Current MCP specifications also include authorization and security building blocks. A large enterprise still has to make its own policy, identity, certification, discovery, and operating choices.

MCP provides The platform adds
A common way for AI applications and servers to communicate Certification, registry ownership, and a production operating model
Standard primitives for tools, resources, and prompts Enterprise identity, policy, and authorization-scoped discovery
Authorization and security building blocks Downstream credential mediation, observability, and audit evidence

These platform responsibilities are PayPal's implementation choices, not requirements imposed by the protocol.

Six questions the protocol leaves to you

  • Which servers are trusted enough to reach production?
  • Which tools may each user discover and invoke?
  • Where do downstream credentials live?
  • How is a human identity preserved through agent delegation?
  • How are sensitive-data controls and rate limits applied?
  • How can hundreds of teams publish without rebuilding the platform?

The first few integrations can answer those questions locally. The next hundred expose the weakness of that model. Local choices become fragmented policy, duplicated credentials, uneven telemetry, and a user experience built from setup instructions.

At this point, connecting more tools was no longer the challenge. Designing a platform to manage them was.

One front door

The Cosmos.AI MCP platform places the PayPal MCP Hub between supported AI clients and certified internal MCP servers. By May 2026, supported clients included Claude, ChatGPT, Perplexity, and VS Code.

They connected to one stable MCP surface, while retaining client-specific capabilities and action-approval experiences. Users reached the certified catalog subject to their own authorization, not a universal list of every server.

The Hub exposes two operations:

  • discover(): find a short list of relevant, authorized tools from a plain-language request.
  • invoke(): route the selected operation through policy to the certified server that owns it.

The Hub at a glance

Conceptual,architecture,diagram:,four,supported,AI,clients,connect,to,a,central,Hub,,which,routes,authorized,calls,to,certified,servers,across,two,trust,zones,while,emitting,operations,and,audit,events.

Conceptual architecture: four supported AI clients connect to a central Hub, which routes authorized calls to certified servers across two trust zones while emitting operations and audit events.

  • Control plane: registration, certification, routing, policy
  • Gateway plane: identity, authorization, limits, routing
  • Certified registry: approved server and tool inventory
  • Observability: operations signals and audit evidence

The client sees a small, stable surface. Behind it, the platform coordinates registration, certification state, identity, policy, routing, downstream credentials, and telemetry.

One request, end to end

Discovery and execution are separate policy events. A plain-language request becomes a tool call in three distinct phases. The Hub limits what can be discovered, the client decides how the action should be presented or approved, and the Hub checks authorization again before execution.

1. Discover (user intent)

  • Ask from a supported AI client
  • Validate identity and caller class
  • Search only the authorized catalog
  • Return a short list of relevant schemas

2. Choose (client boundary)

  • Apply action approval when required

3. Invoke (execution boundary)

  • Revalidate policy and certification
  • Select a separate downstream credential
  • Return the result and emit an audit event

Three boundaries

  • Credential boundary: the inbound Hub token stops at the Hub.
  • Approval boundary: one login is not blanket approval for every action.
  • Attribution boundary: human and executing-agent context stay distinguishable.

One connection does not mean one universal permission. The Hub uses a separately selected, audience-bound credential downstream instead of passing its inbound credential through. Consequential actions can still require client approval, and authorization is evaluated again at invocation time.

The 2,000-tool paradox

A large catalog can make the model less useful. Two thousand tools sound useful until every schema competes for space in the model's working context. In PayPal's 2025 client tests, naively exposing the full discovered catalog to the model produced about 140,000 tokens of tool definitions in the measured 200K-context setup.

That was 70 percent of the available context before the user had asked a question.

Approach (internal 2025 study setup) Tool-schema context
Full catalog test ~140K tokens
On-demand retrieval ~1.6K tokens

That is an estimated 98.9% reduction in this catalog and client setup, returning about 138,400 tokens to the user's work.

The Hub inverted the pattern. Instead of placing the full catalog in model context, it retrieved a small set of relevant schemas on demand. Authorization was applied before retrieval, so the search corpus was already limited to tools the user could invoke.

The finding is historical and environment specific. Current MCP specifications support cacheable list results, and schemas, clients, tokenizers, and cache behavior vary. The durable lesson is not the exact count. It is that a large enterprise catalog needs a deliberate context strategy.

Six design decisions

Make the safe path feel like the easy path. The team organized the platform around six gaps that appeared as MCP moved from pilot to production.

  1. Build and certify as one journey. Maintained scaffolds, shared authentication libraries, automated builds, dependency scanning, registration, and certification made consistent defaults part of the producer experience.
  2. One connection, scoped access. Users connected once to the certified catalog. Identity and policy determined what each person could discover and invoke.
  3. Retrieve tools by intent. People asked for outcomes, not function names. Semantic discovery returned the small set of authorized tools and schemas relevant to the request.
  4. Keep credentials out of clients. The Hub validated its own inbound credential and used separately audience-bound credentials downstream, preserving human and agent attribution.
  5. Match governance to risk. Two trust zones applied deeper review to sensitive systems and an expedited path to eligible lower-risk servers, without creating an ungoverned shortcut.
  6. Let platform gains compound. A new certified capability or retrieval improvement could benefit supported clients through the shared Hub surface, without per-server setup for every user.

From idea to certified capability

Catalog growth also depended on a predictable producer journey. Shared tooling moved common engineering and governance work into a repeatable path:

Scaffold, Implement, Build and scan, Register, Test, Certify, Deploy and observe.

Governance at two velocities

A single review path would either slow low-risk integrations or weaken scrutiny for sensitive systems. The platform used two trust zones.

  • High-sensitivity (restricted-data workloads): deeper review can include threat modeling, dependency analysis, data-flow review, and load testing proportional to the risk.
  • General-purpose (eligible lower-risk workloads): automation plus platform review enabled qualifying servers to complete certification in as little as one business day.

Both paths fed the same certified registry and monitoring model. The rigor changed with the risk. The existence of a gate did not.

Adoption

Growth without a new destination. The Hub launched deliberately with roughly 100 active users in September 2025. A smaller launch protected trust while the team exercised certification, authentication, and scale under realistic use.

Eight months later, in May 2026, the platform served 6,000 active users, including 3,500 monthly active users, across functions that included engineering, product, marketing, support, finance, and legal.

The platform did not ask people to move into another interface. It made their supported AI clients more useful for PayPal-specific tasks. The supply side mattered too: shared builder infrastructure and risk-calibrated certification helped the catalog grow to 175 certified servers in the same period.

The source paper reports an operational case study, not a controlled causal analysis. The figures are historical internal metrics and should be published with PayPal's definitions of active user and monthly active user.

Transferable lessons

Every company's identity, network, and risk model will differ. These five principles travel well.

  1. Treat the gateway as a platform, not a proxy. Routing is one responsibility. Identity, policy, registry, discovery, certification, and observability complete the operating model.
  2. Preserve identity across delegation. An agent action should remain attributable to the person who authorized it and the agent that executed it.
  3. Retrieve schemas on demand. A large catalog is useful only if the model can find the right tools without carrying the entire catalog in context.
  4. Calibrate governance to data sensitivity. Different risk classes need different review depth, but every production path needs an explicit control boundary.
  5. Make the platform disappear for users. One connection, existing identity, and familiar clients turn infrastructure into a capability instead of another destination.

Three takeaways

  • Standardize the connection. Use MCP for interoperability, then make enterprise policy and ownership explicit around it.
  • Authorize twice. Scope discovery before the model selects a tool, then revalidate before execution.
  • Hide plumbing, not accountability. Keep credentials out of clients while preserving approval and human-plus-agent attribution.

The deeper lesson

A protocol became enterprise infrastructure.

MCP gave PayPal a common language. The platform around it made that language operable across identity, policy, discovery, credentials, certification, and telemetry.

Service teams could focus on the systems they own. Users could focus on the task in front of them. One connection, governed end to end, with the complexity carried by the platform instead of the person.


Sources, methodology, and public-claim boundaries

Unless otherwise noted, adoption, catalog, architecture, process, and benchmark figures are based on aggregated PayPal internal telemetry, registry snapshots, and the externally approved case-study paper through May 2026. They are historical point-in-time figures, not public service-level commitments.

The token comparison is an internal estimate from the study's catalog and client setup. Results vary with schema content, client behavior, caching, model, and tokenizer. This article describes Hub-mediated access to certified MCP systems. It does not claim that architecture alone establishes legal or regulatory compliance.

  • Anthropic: Introducing the Model Context Protocol (November 25, 2024)
  • MCP joins the Agentic AI Foundation (December 9, 2025)
  • 2026-07-28 MCP Specification
  • MCP Security Best Practices

Recommended