Cloud Intelligence™Cloud Intelligence™

Read stories from our Forward Deployed Engineers.

Engineering Blog

Trusted by platform engineering teams

Stefanini
Software Projects
Shasta Cloud
Flooid
Finlex
Hippo
Island
Claroty
Salt Security

Latest

All Posts

Your Google API Key Might Be Paying for Someone Else's AI
By John D'EliaMay 11, 20265 min read

Your Google API Key Might Be Paying for Someone Else's AI

Misconfigured Google API keys are being abused to generate AI content on someone else's dime. Here's how the attack works — and a free open source tool to find exposed keys.

You Don’t Have a Pricing Problem. You Have a Cost Visibility Problem.
By Devorah KlartagMay 11, 20265 min read

You Don’t Have a Pricing Problem. You Have a Cost Visibility Problem.

Most SaaS pricing models are built on cost estimates, not cost reality. Learn how runtime cost attribution exposes margin leaks before they kill your unit economics.

You’re Not Ready for the AI Bill That’s Coming
By Devorah KlartagMay 6, 20264 min read

You’re Not Ready for the AI Bill That’s Coming

By the time AI costs show up in your invoice, you're already behind. Here's why getting visibility before you scale matters more than fixing it later.

Top Kubernetes Cost Observability Tools for FinOps
By Devorah KlartagMay 4, 20267 min read

Top Kubernetes Cost Observability Tools for FinOps

The top Kubernetes cost observability tools compared. See which FinOps platforms handle shared resources, multi-tenant attribution, and tagless allocation and how to choose.

Who Foots the Bill? Untangling Google Cloud's API Billing Assignment
By Joshua FoxApr 30, 20269 min read

Who Foots the Bill? Untangling Google Cloud's API Billing Assignment

Google Cloud's API billing model is more nuanced than it appears. Discover how the 'Client' project determines who pays, and how to attribute costs correctly across business units.

LLM gateways are the new FinOps blind spot
By Devorah KlartagApr 29, 20264 min read

LLM gateways are the new FinOps blind spot

LLM gateways collapse every upstream caller into a single workload identity at the provider boundary, breaking cost attribution. Here’s why tagging fails and how runtime (eBPF) attribution reconstructs the chain.

Your AI Writes the Code. Who Owns the Decisions?
By Paul O'BrienApr 27, 202617 min read

Your AI Writes the Code. Who Owns the Decisions?

AI coding tools ship code fast, but leave teams without a record of why. Here's how Kiro's spec-driven workflow and ADRs aim to close that gap.

Instant-On Scaling: Eliminating Node Provisioning Delays in GKE with Active Buffer
By Chimbu ChinnaduraiApr 20, 20266 min read

Instant-On Scaling: Eliminating Node Provisioning Delays in GKE with Active Buffer

GKE's new Active Buffer feature eliminates node provisioning delays by maintaining warm capacity, reducing scale-out latency from minutes to seconds.

The Engineering Guide to Amazon Bedrock Cost Optimization
By Paul O'BrienApr 16, 202616 min read

The Engineering Guide to Amazon Bedrock Cost Optimization

Five proven strategies that compound to 60-80% savings on Amazon Bedrock costs, including batch inference, prompt caching, and model routing.

The AI Coding Paradox: Why More Code Doesn't Mean Better Software
By Birger HalfmeierApr 16, 202617 min read

The AI Coding Paradox: Why More Code Doesn't Mean Better Software

AI coding tools deliver real productivity gains, but don't necessarily improve software quality or organizational learning—a gap worth understanding.

Your cloud bill shouldn't be a mystery

Optimization, automation, expertise. In one platform.

Cloud Cost Optimization Without Tagging: Real-Time Kubernetes Cost Visibility and FinOps Solutions for SaaS
By Devorah KlartagApr 15, 20266 min read

Cloud Cost Optimization Without Tagging: Real-Time Kubernetes Cost Visibility and FinOps Solutions for SaaS

Most FinOps tools can't explain Kubernetes spend by customer or tenant. See how runtime observability replaces tagging for real-time cloud cost attribution.

Stop Node Hunting: How Kubernetes DRA Simplifies GPU Scheduling for AI Workloads
By Chimbu ChinnaduraiApr 7, 20266 min read

Stop Node Hunting: How Kubernetes DRA Simplifies GPU Scheduling for AI Workloads

Kubernetes DRA eliminates manual GPU node hunting by introducing intelligent, request-based allocation for complex AI workloads with mixed hardware requirements.