Trusted by teams where AI inference spend is mission-critical
Connect in minutes
One API key. Full Fireworks visibility.
Connect your Fireworks account with a read-only API key. DoiT ingests billing metrics for serverless inference, on-demand GPU deployments, and fine-tuning jobs automatically. No agents, no code changes, no proxying of inference traffic. You're looking at unified reports within hours of connecting.
What's included
Built for the realities of running on Fireworks
The things FinOps and AI platform leaders actually ask us for when they connect their Fireworks account.
Cost reporting by model
Slice Fireworks spend by model, deployment, fine-tuning job, or team without exporting billing CSVs by hand.
Real-time anomalies
Get alerted on token and GPU spend spikes in minutes, not days.
Deployment recommendations
Know when a workload should move from serverless to a dedicated GPU deployment, and back.
GPU utilization tracking
Spot underused on-demand deployments quietly burning GPU hours.
Fine-tuning cost visibility
Separate fine-tuning jobs and hosted custom models from inference spend, so training costs never hide inside the total.
Governance and budgets
Set budgets per team, model, or environment before spend surprises you.
The billing CSV tells you what you used. Cloud Intelligence™ helps you do something about it.
Beyond the Fireworks billing export
Model and deployment rollups
Consolidated views across every model, deployment, and fine-tuning job, with drilldown into any workload.
Real-time anomaly alerts
Machine-learning detection on model, deployment, and usage-type dimensions, routed to Slack or email.
Spend forecasting
Project token growth and dedicated GPU costs against actual usage before your bill outruns your budget.
Allocation hygiene
Split shared inference spend across teams, products, and environments the way finance expects.
Unified AI infrastructure view
See Fireworks spend alongside Google Cloud and your other providers in a single multi-cloud report.
Forward Deployed Engineers
World-class cloud architects who work as an extension of your team to implement optimizations.
Attribute + Fireworks
Map Fireworks and cloud spend to the work that drives it
Attribute covers Fireworks inference alongside your cloud infrastructure, mapping costs to customers and workloads without relying on a tagging program. Best for: AI product teams that want infrastructure and inference cost attribution in one view.
Serverless inference coverage
Attribute per-token serverless costs to the workloads that generate them.
Fine-tuning attribution
Cover fine-tuning jobs and hosted custom model costs.
Dedicated deployments
Attribute on-demand GPU deployment hours to teams and products.
Full inference API coverage
Capture chat completions, embeddings, and streaming response calls.
Model-level costs
Break down spend across Llama, Qwen, DeepSeek, and your fine-tuned models.
Customer and workload views
Show which customer and workload is responsible for each cost.
Inference anomalies
Detect unusual spend across serverless inference and GPU deployments.
Fast-growing companies run on Cloud Intelligence™
Avg. savings within first 90 days
Avg implementation time
“DoiT's focus on reliability, mixed with the system's flexibility, helps us safely optimize our Amazon EKS workloads with zero-touch from our engineers.”
Oren Ashkenazy
Director of DevOps and Cloud at Fiverr
Ready to rein in your inference spend?
Know what every token really costs.
Frequently asked
questions
How do I get visibility into Fireworks costs across models and deployments?
Connect your Fireworks account once. Cloud Intelligence™ ingests billing metrics for serverless inference, on-demand deployments, and fine-tuning jobs, so you can slice costs by model, deployment, team, or environment from a single view. No manual CSV exports.
What's the best way to integrate Fireworks billing data with Cloud Intelligence™?
Connect with a read-only API key. DoiT handles the rest: ingestion, normalization, and unified reporting across all your Fireworks usage types. Most teams are live within a day.
How can I see which models or workloads drive most of my Fireworks spend?
Cost reports let you drill from total spend down to a specific model, deployment, or fine-tuning job. You can filter by usage type, team, or environment without stitching exports together.
Can I monitor Fireworks cost anomalies in real time?
Yes. Anomaly detection runs continuously across models, deployments, and usage types. When token volume or GPU hours spike, you get a Slack or email alert with the likely cause, long before the invoice arrives.
How do I know whether to use serverless inference or a dedicated deployment?
Cloud Intelligence™ shows your effective cost per workload over time, so you can see when steady token volume makes a dedicated GPU deployment cheaper than serverless, and when a quiet deployment should move back to pay-per-token.
Can I see Fireworks spend alongside my cloud provider bills?
Yes. Fireworks costs sit next to Google Cloud and your other providers in the same reports, budgets, and anomaly detection, so AI inference is part of your total cloud picture rather than a separate silo.
How is Cloud Intelligence™ different from the Fireworks billing dashboard and CLI export?
Fireworks' native tooling shows you raw usage. Cloud Intelligence™ is a platform: unified multi-provider visibility, real-time anomaly detection, budgets and governance, allocation to teams and customers, and access to forward deployed engineers who help you act on what the data shows.
Is my data secure when I connect my Fireworks account?
Cloud Intelligence™ uses read-only credentials with least-privilege access to billing and usage data only. We never touch your models, prompts, or inference traffic, and the platform is SOC 2 Type II certified.
