CommPulse

CommPulse

1159 parked Settings

The cross-site community pulse: gold-layer posts + comment threads read live from the Communication Hub, ranked by importance. Turn a post into Discord / LinkedIn / X.

redditFinOpsimportance 0.60View on Reddit

I am new to FinOps, i wanted to ask people who have experience in the FinOps space. Is there any financial modelling or specifically business case modelling done in FinOps? E.g. if there is any optimisation opportunity or a new workload, is this a requirement from a CFO or board that they need to see a detailed financial model to show ROI and justify the spend? Reason I am asking is because I come from a core finance background just wanted to see if there is an overlap of my finance experience. submitted by /u/Appropriate_Class572 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.60View on Reddit

O365 Outage

by AncientVase

Is anyone else seeing these issues. Just got a call from Help Desk to check it out. Sharepoint home pages are accessible but no files are. Down detector shows a spike but only 128 reports so far. Central/East US region. submitted by /u/AncientVase to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.60View on Reddit

Disclosure: I work with Haevek, a data compute platform. This isn't a pitch and there's nothing to sign up for. Flagging it because this sub asks you to declare. We're putting together a community content series where we interview people who build and run modern data and AI infrastructure. No product talk, no script, no gated landing page. Conversations that get published for whoever finds them useful and nothing gets published without your review. Our view, which you're welcome to tear apart: the bill isn't the software, it's the infrastructure to run it. Open source compute is free to license, but the clusters stay on and consumption pricing climbs with every workload, so cost grows faster than the value coming back. We think the fix is a more efficient engine, not a bigger budget. Plenty of people disagree, which is usually the more interesting conversation. Topics people have picked so far: Where data and AI compute cost actually goes, and why the bill keeps growing as teams do more Scaling AI and agent workloads, where the limit is the cost of running inference over and over rather than the model or the talent What teams get wrong about controlling data and AI cost? Who I'm hoping to talk to: Director or Head of FinOps, Head of Cloud Cost / Cloud Economics VP or Head of Data Platform / Data Engineering who's had the cost conversation forced on them anyone who's actually cut a big data or AI compute line item and can explain what they did Company-wise, anywhere the bill is big enough to be political. Enterprise, scale-up, public sector, doesn't matter. 30 minutes, remote, you get the recording and can cut clips from it. If that's you, or you know someone, DM me. submitted by /u/Sad-Fishing-7666 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditdevopsimportance 0.60View on Reddit

Hi everyone, I would like to hear how experienced DevOps engineers approach monitoring for large public-facing applications. We have a .NET e-commerce platform with: - ASP.NET Core MVC + Angular - SQL Server - Elasticsearch (~10M products) - RabbitMQ - IIS hosting - Multiple public domains/subdomains - Heavy SEO crawling and unknown bots One thing we learned is that monitoring only CPU, memory, and disk is not enough. We have experienced situations where: - CPU and RAM looked normal, but the application was slow - The server was reachable, but users experienced downtime - TCP exhaustion caused issues - Elasticsearch had problems affecting search performance - Bots generated a lot of unnecessary traffic - Slow requests were not obvious from infrastructure metrics I would like to know what metrics and alerts you consider essential for this type of system. Some things I think are important: Application level: - Request rate (RPS) - Response time (p50/p95/p99) - HTTP status codes (4xx/5xx) - Slow endpoints - Exception rate - Thread pool starvation - GC pauses - .NET runtime counters - Memory allocations IIS / Web server: - Current connections - Request queue length - Worker process health - Application pool recycling - Failed requests - Connection errors Network: - TCP connections - TIME_WAIT count - Connection failures - Bandwidth usage - Top clients/IPs - Suspicious user agents Elasticsearch: - Cluster health - JVM memory pressure - Heap usage - Search latency - Query failures - Slow queries - Unassigned shards - Disk usage SQL Server: - CPU - Blocking queries - Deadlocks - Query duration - Connection pool usage - Wait statistics RabbitMQ: - Queue length - Consumer count - Message processing time - Dead letters - Memory usage Security / traffic: - Requests to suspicious paths: - /.env - /.git - wp-admin - Bot traffic percentage - High-frequency clients - Rate limit violations My question: If you were responsible for operating a public .NET application like this, what dashboards and alerts would you consider mandatory? Also, what are some metrics you discovered were extremely valuable only after a production incident? I am especially interested in real-world experience rather than a theoretical checklist. Thanks! submitted by /u/No-Card-2312 to r/devops [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.60View on Reddit

New feature I shipped: --optimize takes the FinOps recommendations CloudCostTree already shows you and applies the ones that are safe to apply mechanically, no architecture or availability trade-off involved. What gets auto-applied: gp2 to gp3 (EBS and RDS storage), provisioned IOPS to gp3 where it's not needed, previous-generation instance types to current-gen, RDS backup retention capped at 30 days, DynamoDB provisioned to on-demand, non-production resources rescheduled to business hours. On Pro with --with-usage, also confirmed orphaned EBS volumes and snapshots, empty target-group load balancers, and unassociated Elastic IPs. What it will never auto-apply: x86 to Graviton (changes CPU architecture), removing Multi-AZ (changes your failover story), CPU or memory based right-sizing from real usage data (measured but still inferred). Those stay as suggestions you confirm yourself. Screenshots show the flow in the VS Code extension against a small test file: the two safe recommendations it found (save $7.01/mo switching off a previous-gen instance, $4.00/mo off a gp2 volume), picking which to apply, and the result, total down from $87.97 to $76.97/mo and the Cost Score up from B to 88. Free tier, and the same thing works from the CLI with cloudcosttree analyze --optimize. https://cloudcosttree.com submitted by /u/Independent-Ease-609 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditawsimportance 0.60View on Reddit

We run a SaaS company in Brazil. Our entire production stack sits in AWS account 393686273302: EKS, RDS, ElastiCache, SES, S3. Application code, customer data, payment records. We think a card change triggered the flag. When we opened the account, the agency that builds our platform (Specter) registered one of their corporate cards so they could provision infrastructure while we set up our own payment method. When the first invoice came due, we replaced that card with our company card and paid the invoice in full. AWS restricted the account after that payment cleared. Our billing console shows R$ 0.00 outstanding today. No open invoice. The suspension notice says non-payment. We have fought verification flags on this account since July: - AWS denied two EC2 vCPU quota increases (L-1216C47A on-demand, L-34B43A08 spot), then granted them after we appealed. - AWS denied SES production access in case 178458595700478, then granted it after we appealed. - CloudFront returned 403 "verification required" and the flag never cleared. - On August 10, RunInstances began returning "This account is currently blocked and not recognized as a valid account". CreateFleet returned MaxFleetCountExceeded while we ran zero fleets. Our Spot quota read 0. Our MediaConvert queue flipped to PAUSED and UpdateQueue returned Forbidden. We opened case 178638951500949 on August 10 and wrote in it that our launch was the next day. We opened a second case; AWS closed it as a duplicate and pointed us back to the first. Nobody from the verification team wrote to us. On August 11st, AWS suspended the account. AWS Health reports our EKS cluster kloel-eks-prod as IMPAIRED: "We couldn't assume the Amazon EKS cluster management service-linked-role" and "We couldn't find or access the AWS KMS key associated with your cluster", with a warning that the control plane shuts down in two days. The NLB in front of our API stopped accepting connections, so api.kloel.com and checkout-api.kloel.com time out from every network we tested. IAM keys that worked before the suspension now return InvalidClientTokenId, so we cannot read our own resources or export a backup. We processed real customer payments hours before the suspension, card and PIX. Those customers now hit a dead platform. We opened case 178647610300640 and uploaded every document AWS requested through their verification link the same day. Nobody has answered us since. Support answered that case with this: > "As this particular inquiry is handled by one of our program support teams, I've forwarded your case directly to them. A member of this team will be in touch with you soon. Our program support team can only communicate through email." We received the same sentence on the earlier case, and nobody contacted us after it. We scheduled our launch for August 11. It did not happen. We are fielding questions from investors and partners about why the product we demoed to them is unreachable. Every minute our platform remains unreachable, the pressure over our company and workers grow, if this goes any further the damage can be inestimable. Two questions for anyone who has been through this: Is there a way to reach a human who can review a payment and reinstate an account? Support forwards our cases and nobody answers. How do we get written confirmation from AWS that our RDS database, EBS volumes and snapshots stay intact while the review runs? We paid the invoice, replaced the card, sent the documents and opened the cases. If anyone from AWS reads this: account 393686273302, cases 178647610300640, 178638951500949 and 178458595700478 have the full history, and we will send anything else you need by DM. submitted by /u/Pandowso to r/aws [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.60View on Reddit

Before deploy: running the analysis against your IaC files up front shows you the full cost breakdown plus savings recommendations, and applies whatever's safe to apply without a human decision, before anything actually gets provisioned. After deploy: once it's live, re-running the same analysis with real CloudWatch usage data (via your own read-only AWS credentials) refines those recommendations against actual utilization instead of static config assumptions. The reason this matters for FinOps specifically: static config tells you what something was provisioned for, not what it's costing you in practice. A right-sizing call made purely from declared instance types will miss real idle capacity, and one made purely from live usage misses waste that never should've been provisioned in the first place. Catching both requires checking at both points in the lifecycle, not just once. (Built this into CloudCostTree, a CLI I've been working on, happy to go into specifics if useful.) submitted by /u/Independent-Ease-609 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

Our CFO caught me after standup and asked a totally fair question, how much is each team spending on all this AI stuff. I said I'd have a number by Friday. Took me until the following Wednesday to admit I couldn't. I'd assumed it would be like AWS where I can slice spend by team in a few clicks, atleast we tagged everything. Then opened the billing expecting the some breakdown but it's just a big monthly number and a graph that goes up. One team lead keeps insisting their usage is not that substancial which without the numbers i cant prove otherwise. We'd handed his squad a shared API key so their spend was all piled onto one. Now, even if I nailed the attribution, most of an agent's bill is the framework re-sending its whole setup every turn, not anything the dev ever typed. I'd be walking into a room to bill someone for tokens they never wrote and can't even see. There's a paper going round with the numbers, arxiv 2607.12161. Anyway. I still owe her that spreadsheet. submitted by /u/Dalius-Gabryelle to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

Doing some research around cloud and AI/token commitment economics and helping NGEN gather feedback on the model. Curious to get the FinOps community’s perspective. The idea is to aggregate compute/token demand across companies, negotiate larger commitments with providers, and use prepayment/financing to offer better pricing and more flexibility. A few things I’m curious about: How much additional savings would make this worthwhile — 5%? 10%+? Is commitment flexibility potentially more valuable than additional savings? Does this make more sense for mid-market companies that don’t already have significant negotiating leverage? NGEN is also collecting anonymous, non-binding indications of demand here (takes ~1 min, no commitment/signature): https://www.ngencompute.com/indication Would genuinely love to hear why you think this would or wouldn’t work. submitted by /u/melc10 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

I’m 22 studying engineering final year. Planning to start out a service business online. I’m looking into providing bookkeeping & FinOps as a service for creative/marketing agencies & eCom businesses Any advise on prerequisites of providing this service from experts in this field could be of help. Thanks, appreciate your time:) submitted by /u/FirstMechanic2151 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.59View on Reddit

Sat through our annual IR tabletop yesterday. Consultant reads a scenario, we all sit around a table, discuss calmly, agree on a plan, write it down, done. Two hours, nobody's blood pressure moved. I've been on the other side of an actual ransomware incident, and nobody was calm. Legal wasn't reachable at 2am, the exec wanted answers before we had them, and half the "plan" from the last tabletop was irrelevant because reality didn't match the script. If your tabletop feels like a calm meeting and not remotely like the real thing, is it actually testing anything? Or is it just a compliance box that happens to have a meeting attached to it? submitted by /u/VegetableFault5149 to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

When AI costs rise and you can't tie that cost to a specific result, project, or team, the safe move becomes turning capabilities off. It makes sense, but it also defeats the point of adopting AI. We just soft-launched Agent Observe in Capital One Slingshot to attribute Snowflake GenAI spend by agent, model, user, etc. using read-only metadata. Early pilot, write-up here . submitted by /u/noasync to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

FOCUS 1.3 solved a real problem: every cloud provider used to bill in its own format, so cross cloud cost comparison meant building custom normalization pipelines just to ask basic questions. Now there’s a common schema. That part is genuinely good. But adoption of the schema is not the same as fixing FinOps. I keep seeing teams roll out FOCUS, get their data normalized, and still can’t answer the question that actually matters to leadership: who owns this cost, and why did it move. The reason is that FOCUS standardizes the shape of usage and cost data. It says nothing about your tagging discipline, your allocation model, or who is accountable when an engineering team spins up something that triples a bill overnight. Those three things are where the actual FinOps work lives, and they’re organizational problems, not schema problems. You can have perfectly FOCUS compliant data and still have zero cost accountability, because accountability comes from tagging governance and process, not from the spec. The teams that get real value from FOCUS are the ones who treat it as the foundation for building an allocation and accountability model, not as the finish line. If your rollout stopped at “we ingest FOCUS data now,” you’ve done the easy 20 percent. I went deep enough on this that I ended up writing a book on it, “Cloud Money,” working through the FOCUS spec at a practitioner level alongside cost allocation strategy and accountability models. Not trying to sell it here, just flagging it in case anyone wants the longer version of this argument. Happy to talk through the tagging governance side in the comments if people have specific setups they’re stuck on. submitted by /u/aliihashmiii to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.59View on Reddit

I am working on Azure Data Explorer to optimise its cost usage for my clients, currently I am suggesting downsizing the engine instance SKU type based on the following metrics : CPU utilization < 45% Cache utilization factor < 55% IngestionUtilization and StreamingIngestUtilization < 45% If these holds downsize to the very next lower sku type. But the blocker i am having is that , there's this limit of 50% ram/node for a query , so this will break and cause my queries to fail , is this for real or only latency will increase. Can anyone help me out with that. submitted by /u/sirius_black19 to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

I've been building Cognocient for a few months now, a proxy that sits between your app and OpenAI/Anthropic/Gemini and tells you what each feature actually costs, before the call goes out instead of after you get the bill. For a while the only way in was a 10 day trial with full access, then you had to pick a paid plan. That made sense for teams evaluating it for real budgets. It made no sense for someone who just wants to drop it into a side project and see what their AI calls actually cost. So I added a free tier that doesn't expire. What's in it: one provider connection, the real time proxy and attribution dashboard by feature and model, one budget with alert level enforcement, 7 day retention. Capped at $50/mo of tracked spend, after that the proxy keeps forwarding your calls (I will not break your app over a free tier limit) but stops logging new attribution until the next cycle. Setup is one line, you swap your base\_url for the Cognocient proxy URL and nothing else in your code changes. If you're already tracking spend some other way I'd genuinely like to know what you're using, half the reason I built this is I couldn't find anything that did pre-call enforcement instead of after the fact dashboards. [cognocient.com]( http://cognocient.com ) if you want to poke at it. submitted by /u/MaverikSh to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

New Harness - costs per run

by imnotonetogossipbut1

Does anyone have any direct experience of the typical costs coming in for the new harness ? Microsoft are making a number of general statements about "long running, multi-step jobs" but still no clear breakdowns of what a particular businss workflow (e.g. receive email from customer, look up record in salesforce, determine sentiment, notify account manager if needed, arrange customer manager meeting, send confirmation". might cost. Because testing is now no longer free, its hard to run through these models and look at typical charge cards without some good examples. Under the old model we had clear credit values for doing certain actions. submitted by /u/imnotonetogossipbut1 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X