Setup
One API key. Full gateway visibility.
Connect your Requesty account with a read-only API key. DoiT ingests spend, token usage, and per-model costs from your gateway automatically. No agents, no code changes, no re-instrumenting your LLM calls. You're looking at unified reports within hours of connecting.
What you get
Built for the realities of running LLMs through a gateway
The things FinOps and engineering leaders actually ask us for when they connect their Requesty environment.
Unified token cost reporting
Slice Requesty spend by model, provider, API key, or team without exporting usage logs into spreadsheets.
Real-time anomalies
Get alerted on token usage spikes in minutes, not hours.
Per-model cost breakdown
Compare what Claude, GPT, and Gemini requests actually cost you side by side.
Provider spend rollups
Track spend across every provider your gateway routes to in one view.
AI and cloud spend together
Put Requesty gateway costs next to the cloud infrastructure serving your app, so total unit costs finally add up.
Governance and budgets
Set budgets per team or product without chasing down every API key.
Requesty's dashboards tell you what your gateway spent. Cloud Intelligence™ helps you do something about it.
Beyond Requesty's built-in analytics
Cross-provider rollups
Consolidated views across every model provider Requesty routes to, with drilldown into any model or key.
Real-time anomaly alerts
Machine-learning detection on model, provider, and team dimensions, routed to Slack or email.
Spend forecasting
Project token spend against actual usage trends so AI budgets don't surprise finance at month end.
Allocation hygiene
Attribute LLM spend to products, features, and teams, and split shared costs the way finance expects.
Full-stack cost context
See gateway spend alongside the Kubernetes and cloud costs serving the same application.
Forward Deployed Engineers
World-class cloud architects who work as an extension of your team to implement optimizations.
Attribute + Requesty
Map Requesty and LLM spend to the work that drives it
Attribute maps AI inference costs to customers and workloads without relying on a tagging program. Combined with Requesty's gateway, every routed request becomes attributable spend. Best for: AI product teams that want gateway and infrastructure cost attribution in one view.
Gateway request coverage
Attribute costs across every request routed through Requesty.
Provider attribution
Cover OpenAI, Anthropic, Google, Mistral, and other providers.
Serving infrastructure
Attribute the compute behind your app alongside gateway spend.
Full LLM API coverage
Capture completion, streaming, and tool-call requests.
Per-model costs
Break down spend across Claude, GPT, Gemini, and other models.
Customer and feature views
Show which customer and feature is responsible for each token cost.
AI spend anomalies
Detect unusual token usage across models and providers.
Fast-growing companies run on Cloud Intelligence™
Avg. savings within first 90 days
Avg implementation time
“DoiT's focus on reliability, mixed with the system's flexibility, helps us safely optimize our Amazon EKS workloads with zero-touch from our engineers.”
Oren Ashkenazy
Director of DevOps and Cloud at Fiverr
Ready to connect your Requesty gateway?
Bring clarity to your LLM token spend.
Frequently asked
questions
Is there a single view of LLM costs across all the providers Requesty routes to?
Connect Requesty once. Cloud Intelligence™ ingests spend and token usage for every provider and model your gateway routes to, so you can slice costs by model, API key, team, or product from a single view. No exports, no manual rollups.
What does connecting Requesty to Cloud Intelligence™ involve?
Use a read-only API key from your Requesty account. DoiT handles the rest: ingestion, normalization, and granular per-model reporting. Most teams are live within a day.
Which models or teams account for the bulk of my token spend?
Reports let you drill from top-level gateway spend down to a specific model, API key, or team. You can filter by provider, model, or time window without writing queries or parsing usage logs.
Will I be alerted to token usage anomalies as they happen?
Anomaly detection runs continuously across models, providers, and team dimensions. When a runaway loop or a misrouted workload spikes your token usage, you get a Slack or email alert with the likely cause before the spend becomes a real problem.
Can I set budgets on Requesty spend per team or product?
Yes. Set budgets and alert thresholds per team, product, or API key in Cloud Intelligence™, and get notified as spend trends toward the limit rather than after the invoice lands.
I already use Requesty's built-in analytics. What does Cloud Intelligence™ add?
Requesty's dashboards show real-time gateway usage. Cloud Intelligence™ is a platform: it puts that spend next to your cloud infrastructure costs, adds anomaly detection, forecasting, budgets, cost allocation, and access to forward deployed engineers who help you act on what the data shows.
Can I see Requesty costs alongside my cloud infrastructure spend?
Yes. Cloud Intelligence™ unifies LLM gateway spend with your cloud bills, so you can see the true cost of an AI feature: model inference plus the compute, storage, and networking that serve it.
What access does Cloud Intelligence™ get to my Requesty account, and how is it protected?
Cloud Intelligence™ uses a read-only API key with least-privilege access to usage and billing data. We never touch your routing configuration or make changes without your approval, and the platform is SOC 2 Type II certified.
