ZeroGPU × ZeroClick: specialized inference any agent can buy

Frontier models are spectacular generalists, but there’s a great deal of production AI work that can be done using more efficient models. ZeroGPU built an inference network on that observation - specialized small and nano language models, running on edge compute, at a fraction of the frontier cost. Today we're announcing a partnership that makes ZeroGPU's entire model catalog purchasable by AI agents: per call, over open payment rails, with no API key in sight.
The compute-efficient layer for AI inference
ZeroGPU routes high-volume AI tasks - classification, extraction, summarization, PII redaction, content moderation, and more - to specialized small language models across an edge-powered inference network. The company's line is "use frontier models for reasoning, use ZeroGPU for everything else," and its published numbers make the case: 50%+ lower cost on routine workloads and up to 10x faster results on classification and signal extraction. In ZeroGPU's own case study with a partner, moving adtech tasks like IAB classification from frontier LLMs to its small language models cut AI spend by more than 6x while cutting latency 10x.

The APIs are OpenAI-compatible - and with integrations for OpenClaw, Claude, and MCP support, anything that can call a chat completions endpoint can use ZeroGPU. Now, with this partnership, ZeroGPU is ready to do business with agents.
What the partnership does
Every model in ZeroGPU's catalog is now its own agent-purchasable service on ZeroClick, metered on input and output tokens, priced per call. An agent discovers the storefront, reads the machine-readable listing, pays over x402 or MPP, and gets its completion back. Revenue settles directly to ZeroGPU's Stripe account.
Our favorite detail: there are no API keys in the flow. ZeroClick verifies the agent, handles payment, and forwards each request with a signed header that ZeroGPU's API checks.
By setting up an agent storefront, ZeroGPU opened a new sales channel without having to change its authentication model or issue credentials to strangers.
Discovery comes built in. ZeroClick automatically registers the agent storefront across major agent indexes - Coinbase's x402 discovery platform agentic.market, x402scan, mppscan, and zero.xyz - so ZeroGPU's models show up wherever agents search for inference.
Since launching their agent storefront, ZeroGPU has added popular open-weight models like DeepSeek, GLM, Kimi and more because agent demand called for them. That's a new kind of signal for a product team: not a sales forecast, a transaction log.
"ZeroClick built the storefront for us. We connected Stripe, verified the services, and started selling to a customer segment - agents - we couldn't easily reach before."
- Maddy Arvapally, ZeroGPU
When the buyer is an agent, efficiency wins
An agent doesn't have a favorite model. It has a task, a budget, and a latency target. That makes agents the most rational buyers specialized inference has ever had: an agent that needs ten thousand log lines classified doesn't want a frontier generalist, it wants the cheapest model that clears the quality bar - and it can read the price and the benchmark before it pays. ZeroGPU wins that comparison on exactly the workloads it was built for.
Get started
If your workloads don't need a frontier model, try ZeroGPU - your agent can pay for exactly what it uses, straight from the storefront. And if you're ready to sell your own product to AI agents, that's exactly why we built ZeroClick - get a demo at zeroclick.ai, or read the docs at docs.zeroclick.ai.