Karpenter launched 22 × m6i.8xlarge in 4 minutes
PR #1832 · matchmaking · game-services · role karpenter-prod
Real-time cloud cost control for AWS and Kubernetes
Cloudblame turns CloudTrail, CloudWatch and Kubernetes signals into named findings: the deploy, alert rule or service account behind a cost spike, a reversible brake, and savings verified on your own bill.
Connects in minutes. Read-only role with your external ID, nothing to install.
FEED · FROM THE DOCUMENTED EKS AND ATHENA CASES
Reads from 14 sources. Writes to none.
CLOUD COST ENGINEERING PLATFORM
Cloudblame reads your usage signals as they happen, names the change behind a spike, hands you a reversible brake and verifies the result on your own bill. One finding id carries through all four steps.
READ-ONLY BY DESIGN
The exact policy and the CloudFormation template are public: policy.json · readonly-role.yaml
PROBLEM
A merged pull request, an alert rule or a service account can move your AWS bill within minutes. Cost Explorer and anomaly alerts report it a day later as service, account and usage type, and leave the team to find the change by hand.
Cloudblame closes the loop between usage signals and a verified fix.
MANUAL INTERVENTION REQUIRED. CLICK PODS TO BRAKE.
THE CORRELATOR
Six links from a line on your bill to the pull request and the team behind it. AWS Cost Anomaly Detection stops at the account and the IAM principal; Cloudblame keeps walking the chain until it has a name.
AWS tells you which role launched the nodes. Cloudblame follows the NodeClaim to the pending pods, the Argo CD revision and the pull request that merged at 14:02.
Savings are measured on your Cost and Usage Report against a unit price frozen at the finding. Later price changes and discounts do not move the number.
Every brake is a change you run in your own account and can undo: revert the PR, cap the node pool, widen the alert interval.
CASE FILE
A pull request raised CPU requests from 1 to 8 vCPU on 40 replicas; Karpenter added 22 nodes in four minutes. AWS reported it the next day as “EC2 usage increased”. Cloudblame named the PR, the team and a one-line fix at 14:09.
EKS · MATCHMAKING · PR #1832 · $24,300/MONTH VERIFIED
Red: Cloudblame. Grey: AWS, the next day. Green: verified on your Cost and Usage Report.
22 × m6i.8xlarge launched for game-services/matchmaking in 4 minutes
HOW IT WORKS
Launch one CloudFormation stack: a single read-only role, your external ID. About ten minutes, nothing installed in your cluster.
~10 min · read-onlyA free one-page answer: what AWS already explains, what it cannot, and a clear go or no-go before you spend another hour.
free · one pageNamed findings with apply-ready fixes and a reversible brake. You approve every change; nothing runs without you.
you approve every changeThe delta is measured on your Cost and Usage Report against a frozen baseline. You pay only on that number.
measured on your CUR| Key | Value | Description |
|---|---|---|
| RoleArn | arn:aws:iam::123456789012:role/CloudCostControlReadOnly | Paste into step 2 of the scan |
| ExternalId | 7f3a19c4…e91b | Generated in your browser |
| PolicyActions | 55 | ce, cloudtrail, cloudwatch, athena, ec2, eks · read-only |
| CurQueries | false | Switch on later for the measurement annex |
INTEGRATIONS
Every signal comes from services you already pay for. Findings land where your team already looks: a Slack card and a comment on the pull request that caused the spike.
READS · ONE READ-ONLY ROLE, ~55 ACTIONS
AWS · the published role
cloudtrail:LookupEvents cloudwatch:GetMetricData · logs:StartQuery ce:GetCostAndUsage · ce:GetAnomalies athena:StartQueryExecution · s3:GetObject on the CUR bucket ec2:DescribeInstances · ec2:DescribeNatGateways eks:DescribeCluster · eks:ListNodegroups cost-optimization-hub:ListRecommendations Around AWS · read-only tokens
get, list on pods, deployments, nodeclaims · view ClusterRole applications: get contents: read · pull requests: read alert rules: read · usage: read Analytics: Read · Cache Rules: Read The exact list is public: policy.json · readonly-role.yaml
WRITES BACK · TWO PLACES, NOTHING ELSE
22 × m6i.8xlarge launched for game-services/matchmaking in 4 minutes
This PR raised matchmaking CPU requests 1 → 8 vCPU on 40 replicas. Karpenter added 22 nodes (704 vCPU) to node pool general at 14:03; p95 usage is 0.4 vCPU per pod.
Suggested fix: requests back to 1 vCPU, HPA at 70% → PR #1840
READ-ONLY ROLE · ~55 ACTIONS, PUBLISHED POLICY · NOTHING INSTALLED IN THE CLUSTER
PRICING
Net means after any cost the fix introduced, like a VPC endpoint. Verified means measured on your Cost and Usage Report against a frozen baseline. Recommendations are never invoiced.
Read the measurement annex| Line | Fee |
|---|---|
| Commitments (RI / Savings Plans) | 0% |
| First $10k / month | 10% |
| $10–30k / month | 7% |
| $30–60k / month | 5% |
| Above $60k / month | 4% |
| Per finding | 12 months, then $0 |
| Under $500 / month | not billed |
| Diagnostic, platform, setup fees | none |
A 20%-of-annual consultancy would take for the same savings.
| Finding | Baseline | Actual | Net saving | Fee |
|---|---|---|---|---|
| #003 EKS matchmaking requests 8 → 1 vCPU (PR #1832) Verified | 704 vCPU | 64 vCPU | $24,300 | |
| #005 NAT → S3 gateway endpoint (payments) Illustrative | 2.1 TB/day | 0.3 TB/day | $4,860 | |
| #007 DynamoDB orders table, provisioned → on-demand Illustrative | 3,000 WCU | 640 WCU | $2,180 | |
| Net savings in this example | $31,340 | |||
| Fee · 10% of first $10k + 7% of next $20k + 5% of $1,340 | $2,467 | |||
PLAYBOOKS
Each card is a playbook: the signal we watch, how fast it fires, where AWS stops and what we add. The same finding id follows it from the Slack card to the month-end statement.
AWS seesService, usage type, principal
We addThe alert rule or dashboard, its owner, the fix
AWS seesUsage type
We addThe log group, the deploy that raised the level, who queries it
AWS seesAn idle NAT
We addThe pod and the missing VPC endpoint
AWS seesThe list (free)
We addWho created it, when, and the PR to remove it
AWS seesNode level
We addNamespace, workload, owner, the requests to change
AWS seesThe principal
We addThe retry setting and the loop
AWS seesNothing on data-plane calls
We addThe requester and the lifecycle rule
AWS seesA day later
We addPriced in minutes, with a brake
AWS seesBedrock principal
We addThe loop, the prompt growth, the model choice
AWS seesNothing
We addThe cache rule that sends bytes to origin
AWS seesNothing
We addThe deploy that added the cardinality
FREE SCAN · 10 MINUTES · READ-ONLY