Cloud Intelligence™Cloud Intelligence™

Announcement

This page is also available in Deutsch, Español, Français, Italiano, 日本語, and Português.

See what every model behind your Requesty gateway really costs

Connect your Requesty organization to Cloud Intelligence™ and see LLM spend per model, hosting provider, API key and routing policy, priced at what you pay for credits, next to the rest of your cloud spend

An LLM gateway is useful precisely because it hides things from your application: which provider served the request, which region, which fallback fired when the primary hit a rate limit. That is what you want in the request path and not what you want on the invoice. Requesty draws each request from a prepaid credit balance at the provider's list price and charges its 5% when you top up, so the per-request figures in its dashboard sit 5% under what finance actually paid. Usage is keyed by model name, so gpt-6-astra served from OpenAI and gpt-6-astra served from Azure land under one label unless you also split by provider. Working out what one team spent on Claude via Bedrock versus Vertex last month means an export, a spreadsheet and a credit-card statement.

Cloud Intelligence™ now reads usage straight from your Requesty organization and prices every row the way your invoice does. The report that shows what your EKS cluster cost yesterday can show what it spent on tokens through Requesty, down to the API key and the routing policy that picked the model.

What you get

Daily spend and token usage for every model your Requesty organization touched, imported into Cloud Analytics as a first-class provider. The SKU is hosting provider plus model, so openai/gpt-6-astra and azure/gpt-6-astra stay separate and bedrock/claude-haiku-4-5@us-east-1 is its own line. Labels give you the hosting provider, the API key, the Requesty organization, and the routing policy when a fallback, load-balancing or latency policy chose the model rather than your code. Input tokens, output tokens and request counts come along as metrics.

The numbers mean what finance thinks they mean. Cost is the credits drawn plus Requesty's 5% markup, so it matches the credits you buy. List cost is the provider list price, so it matches the Requesty analytics dashboard. On a test organization we reconciled both to the cent against the balance debit. Bring-your-own-key requests are billed by the upstream provider directly, so they carry zero cost here to avoid double counting your OpenAI or AWS feed; their estimated provider spend is kept as a separate metric, genai.cost.byok_external.

Because Requesty lands as a regular cost and usage source, everything downstream just works: reports, budgets, anomaly detection, forecasts, and the GenAI Intelligence dashboard, where Requesty shows up next to OpenRouter, Anthropic, OpenAI, Fireworks AI and your other inference spend.

Setup is deliberately boring. You create one API key in the Requesty Console with Manage permission set to Read and Completions set to None, pick a name for the connection, and you're done. Put a small monthly spend limit on that key as well; Cloud Intelligence™ only reads usage and never sends inference requests. On first connect we backfill up to 12 months of history, then refresh every 6 hours as Requesty closes each UTC day.

Get started

  1. Create a management API key in the Requesty Console with Manage: Read, Completions: None, and a monthly spend limit.
  2. Connect your Requesty organization in Cloud Intelligence™ under Data ingestion and integrations > Integrations, and run Test connection.
  3. Build a report grouped by SKU, or by the Provider, API key and Routing policy labels, once the import email arrives. The Requesty reporting dimensions list what maps where.
  4. Open the GenAI Intelligence dashboard to see Requesty alongside your other GenAI providers.

The Requesty connector is [CONFIRM: available now to all Cloud Intelligence™ customers] and is the fourth LLM gateway we support, after OpenRouter, LiteLLM and Bifrost. If you run Requesty in production, connect it and tell us where the numbers disagree with your invoice. Bring-your-own-key traffic in particular has only seen our test organization so far, and real volume from a real org is what we need next. Your account manager or an expert inquiry is the fastest route.

PerfectScale™ for Kubernetes

Ready to optimize?

Get your free Kubernetes savings analysis