CommPulse

CommPulse

1160 parked Settings

The cross-site community pulse: gold-layer posts + comment threads read live from the Communication Hub, ranked by importance. Turn a post into Discord / LinkedIn / X.

redditdevopsimportance 0.59View on Reddit

Hey hi everyone, I am just trying to understand how engineers/SREs who dealt with real production latency incidents investigate it Lets say you have the following - Logs - Recent deployment information - Application health - Database metrics - External dependency health/metrics - Infrastructure metrics You just encountered the incident, you dont know the root cause. You are uncertain about the truth. From here how do real engineers go about reasoning to find the root cause - Do you follow a standard sequence of investigative steps - How do you determine what investigative step to take next under uncertainty to narrow down the possibilities for the root cause - Have u ever encountered with incident where initial information was misleading, how did you navigate from there - Is there any situation where you have lot of information but struggled to form a proper hypothesis - Have you tried any AI investigative tools that help you in achieving this I just wanted to understand how do real engineers reason through the uncertainty to find the root cause. What are the biggest pain points submitted by /u/HawtBeagle to r/devops [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

An 80 percent price cut makes a nice chart. It does not make an 80 percent cheaper task. LLM cost optimization starts with the task ledger, not the price sheet. The useful unit is an accepted task: work the team is willing to ship. Each row needs model tier, uncached input, cache traffic, output, retries, review time, and a final accepted or rejected flag. I keep the accounting wrapper fixed by routing Sol, Terra, and Luna through ZenMux's multi-model API gateway and changing only the model slug at one endpoint. That still leaves provider behavior, workload, acceptance rate, and retries. A cheap model can get expensive when review or reruns creep in. Until the ledger has those rows, the July 30 headline is just a new line in the price sheet. submitted by /u/Dramatic_Spirit_8436 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

I assume everyone has been following the AI build out. Memory prices have risen sharply. GPUs, CPUs, memory shortage is hitting the consumer market. Apple announced price increases in hardware, which is quite rare for Apple. We run on Hetzner and we reserved several Hetzner instances a few months back. The renewal prices have risen since we reserved them. We also run on AWS, but in much smaller numbers. We haven't seen any major changes in our AWS bills thus far. But for folks who are operating much larger accounts, I'm trying to figure out is when these will hit AWS/GCP/Azure list prices. I personally think its a matter of "when", as opposed to "if". If you renewed a savings plan, RI or CUD recently, was the effective rate worse than the term it replaced. Has an account team given anyone a heads-up? I work at Readyset which is a caching solution for databases, so we have an obvious interest where instance costs go. I'm trying to work out whether this is a real 2026 budget line or mostly bare metal hosts that have less pricing cushion than the hyperscalers. submitted by /u/Master-Bass-1905 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

looking at a pretty big gpu bill right now. the normal cloud cost tools can obviously tell me what the instances cost. but im trying to get more granular. cost by workload. team. job. maybe even gpu utilization vs what were actually paying for. are there any finops platforms that do this well? what are you guys using? submitted by /u/eastwill54 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

I am hoping to get an outsiders perspective, and thank anyone in advance for taking the time to respond. I feel like there could be other young professionals in my situation that this will hopefully help. Background Non-technical, non-accounting degree with further analytics qualifications UK based with ~5 years of experience both as a consultant and in-house and have actioned millions in savings initiatives working with engineers etc. I am the first FinOps hire that my current company has ever made. Exposed to multiple clouds (private and public), industries and technologies. A range of vendor certifications (AWS, Azure etc.) I genuinely love the work I do, it is the perfect intersection between technology and finance that scratches a very specific itch. But given my background I feel like I can see an upper limit in terms of career trajectory and am wondering which path I should take and what are the steps required to avoid getting stuck as an analyst. Problem I can talk the talk with engineers but I do not currently possess the technical skills to go down an engineering/architecture path, although "cloud architecture" is probably my favourite part of the job even if I feel like I am just successfully guessing most of the time. The imposter syndrome that comes with that is less than ideal. My lack of accounting/finance background makes me think that I would struggle to be taken seriously in a more senior role. I would be curious to get your guys' take on this: Am I just overthinking and experience will make up for my slightly non-conventional background? Do you have any recommendations for financial or technical qualifications (outside of the usual vendor certs)? This is a multi-year strategy. Given my lack of engineering background, am I better off leaning into finance/accounting? Cheers! submitted by /u/EconThrowA to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

UBS put out numbers showing the big three cloud providers are spending about 102% of their cloud revenue on capex, $4.1 trillion is projected through 2028. that spend lands somewhere, and what i'm watching is whether it hits non-AI workloads. OVHcloud already raised prices citing the memory shortage, some servers up as much as 87%, because AI is eating the same memory and power regular nodes run on. the thing is it won't show up as a line item. it's compute and memory quietly drifting up with no change in usage, and if your cost tooling only watches your own consumption, it won't flag a provider-side price move. so you can't really tell whether your usage went up or their prices did. anyone renewed an RI or savings plan lately and had the rate come back worse than the term it replaced? submitted by /u/CryOwn50 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

Declared self-promotion per the sidebar: my project, launched this week. https://llmcostkit.com Free, client-side calculator answering the two questions FinOps keeps getting asked about AI spend. First, what do LLM API tokens cost against a real workload: 24 current models across OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, Meta-hosted Llama and Perplexity, with prompt caching and batch discount modelling, projected from your requests/day and token sizes. Second, what do the per-seat tools cost: M365 Copilot, GitHub Copilot, Cursor, Kiro, ChatGPT Business and Claude Team, against your seat count. Prices were verified against vendor pages on 12 Aug 2026 and the page says so on every table, because half the calculators out there quietly serve stale numbers. Filter to the models you actually use and it gives you one answer line. CSV export, shareable scenario links. No signup, no tracking, no cookies. Installs as a web app and works offline. One deliberate omission, explained on the page: Databricks Mosaic AI, because DBU-based pricing has no honest single number, so there's a custom-rate row instead of a made-up one. Transparency: the site sells a paid kit for teams (unit economics model, allocation and tagging policy templates, CUR/Azure/BigQuery starter queries, chargeback model, maturity assessment). The calculator is free regardless and doesn't nag you about it. What's missing that would make this genuinely useful for your practice? PTU and provisioned-throughput break-even modelling and self-hosted GPU comparison are top of my list. I'll build the most-asked. submitted by /u/TeacherInevitable408 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

Clodkeeper AZ?

by DReddit111

Our AWS rep hooked us up with CloudKeeper. They pitched us their AZ product and said we could get a 2% discount off our AWS bill and free support, plus some finops tools, the product doesn’t cost us anything, they don’t have to take over our AWS root account and we can leave with 60 days notice. Said they make their money on bulk AWS discounts that get that they share with us. It’s not a huge amount of savings, but there doesn’t seem to be a downside. Anybody have any experience with them. I’ve been in the business for a long time and there is usually a gotcha in there somewhere. submitted by /u/DReddit111 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.59View on Reddit

Every FinOps conversation about AI cost I have run into loops back to cost per token. It is the number vendors publish, so it feels concrete. It is also the wrong number to argue about. At the AI deployments I have worked on close enough to see the real numbers, the token bill was rarely more than a third of the actual TCO. The rest sat in three places nobody was tracking as tightly. GPU underutilization at inference is the first one. Reserved capacity sitting at single-digit average utilization is normal, not exceptional. Teams blame batching. The real cause is a prompt-mix distribution nobody profiled before signing the reservation, and the invoice for that gap does not carry a "token" label. Storage is the second. Vector stores, eval traces, and audit logs outpace the token bill within a couple of months of any real RAG workload going live. It is not that any single thing is expensive. It is that nobody set a lifecycle policy at design time and the growth curve is invisible until it is not. Governance is the third and the most awkward, because most FinOps units skip it entirely. Evaluation pipelines, red-team runs, human-review loops, policy scans. Engineering time and pipeline compute, not a line on the AI vendor invoice, but it is TCO. Anyone who runs a compliance-adjacent workload has felt this bucket outgrow the token bucket without ever showing up on a cost dashboard. The docs and pricing pages train us to argue about fifteen cents versus thirty cents per million tokens as if that is the FinOps decision surface. It is the marketing surface. So the practitioner question. What unit does your team actually use for AI workloads? - cost per token - cost per successful task or workflow - cost per active user per month - cost per business outcome (ticket resolved, fraud caught, revenue attributed) Or is your team stuck between the vendor unit and the business unit with nothing that stays honest under load? submitted by /u/matiascoca [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.57View on Reddit

Hey everyone, Over the past few months, I’ve been analyzing enterprise AI billing data and studying why so many engineering teams and companies are getting hit with massive, un-modeled AI invoices. For two years, the industry narrative has been that AI is getting dirt cheap and price per token keeps dropping exponentially. Yet, across Big Tech and mid-sized companies alike, actual monthly invoices are skyrocketing. Here is a quick breakdown of the mechanics behind why this is happening: 1. The 1865 Jevons Paradox is alive in Tech In 1865, economist William Stanley Jevons observed that when steam engines became dramatically more efficient at burning coal, Britain didn't burn less coal, it burned exponentially more. Why? Because cheap coal suddenly made financial sense in places where nobody could justify the cost before. The exact same thing is happening with LLM tokens. As unit costs drop, consumption doesn't stabilize but it expands into every workflow, background agent, and automated task until nobody weighs the unit cost anymore. 2. Real-world corporate overruns Uber: Handed a coding agent to 5,000 engineers. By April, just four months into a 12-month plan, their entire annual AI budget was completely gone. The tool was so useful that usage exploded. Meta: Built an internal leaderboard ranking engineers by token burn rate. In one month, they burned 73.7 trillion tokens before executives realized token burn measured activity, not actual impact, and killed the board. Microsoft: Ordered internal divisions off external coding tools days before their fiscal year closed to force migration onto cheaper internal alternatives. 3. The agent multiplication factor (5x - 30x Tokens) Standard chatbots are 1-input / 1-output. AI agents are fundamentally different. Because current architectures lack long-term memory, at every loop step (plan, search, tool call, handoff), an agent must package the entire conversation history and re-submit it to the API. Data from Gartner shows an AI agent burns 5 to 30 times more tokens than a basic chatbot doing the exact same task. Token prices dropped 60%, but agent loop usage increased 1,000%. 4. The hidden "Second Meter" Every time an agent writes a code block or report and a human engineer spends 30 minutes reading, verifying, or rewriting it, you pay twice: once in API tokens, and once in senior engineering salary. I put together a full 17-minute video essay breakdown with all the diagrams, data sources, and frameworks (including OpenAI CFO Sarah Friar’s scorecard on measuring "useful intelligence per dollar") here: Watch the full breakdown here: https://www.youtube.com/watch?v=DBf5-yBRxEk Curious to hear from engineering leads, FinOps folks, and founders here: How are your teams tracking agent loops and token spend right now? Are you capping per-user usage, or waiting for the quarterly invoice to arrive? submitted by /u/thadah123 [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.52View on Reddit

https://preview.redd.it/d1obi7fs6ylh1.png?width=389&format=png&auto=webp&s=a600d034a9db5722d606b1e4ba102d9af27d7e17 https://preview.redd.it/84eonkfy6ylh1.png?width=1564&format=png&auto=webp&s=48f9e421ecc1acff9d67b59a62f4bae3c3ba5871 https://preview.redd.it/602423b37ylh1.png?width=1858&format=png&auto=webp&s=2fb26bee0fac2afb2ef94535a36b0be2781b3648 I'm having a strange intermittent issue with an Azure VM and would appreciate help troubleshooting it. VM details: Standard D4ds v4 4 vCPUs 16 GiB RAM Ubuntu 24.04 Static public IP DNS managed through GoDaddy Frontend and backend running on the same VM The VM doesn't completely stop responding. The behavior is: Normally SSH is fast. During the incident, SSH still works, but takes around 3–4 seconds to connect . Once connected, typing commands has around 1 second of noticeable delay . At the same time, the website goes completely down/unreachable for some period. The frontend/backend processes don't appear to crash. After the issue clears, SSH becomes normal again and the website starts working again. I don't intentionally restart the applications when this happens. Azure's Availability metric also shows several drops around these periods. I've attached the screenshot. The domain is managed through GoDaddy and points to a static public IP , so the IP isn't supposed to change. What I'm trying to understand is what could cause this combination: SSH = extremely slow/laggy Website = completely unreachable Applications = still running Azure Availability = intermittent drops Could this indicate packet loss/network problems, an Azure NIC/public IP issue, VM resource contention, or something happening at the Ubuntu networking level? What should I monitor during the next incident to determine exactly where the problem is? Thanks! submitted by /u/Ok_Cat_2052 to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.52View on Reddit

This was brutal, took down out production app for a bit, couldn't connect to the keyvault. Here's a link to the incident if others saw this issue. Every time something like this happens, I wonder if this type of this is going to happen again. I suppose it's fixed now, but I'm asking support for a full report. And we're on premium as well. https://app.azure.com/h/9Y7C-GKZ/9cf095 submitted by /u/dptech3 to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditawsimportance 0.52View on Reddit

Aws locked us out

by Zealousideal_End_366

Small startup locked out of our AWS account with no explanation — support has stalled for half a day. Can anyone help escalate? We’re a small startup and our entire AWS account got locked earlier today with no warning and no clear reason given. Our production environment is down and we can’t access anything. We opened a support ticket right away, but it’s now been more than half a day. It’s bounced from the service team to the security team and back again, with no resolution and very little communication in between. Every hour of downtime is a serious hit for a company our size. I’m not looking to bash anyone — I just need this in front of someone who can actually move it forward. If there’s an AWS employee, an MVP, or anyone who’s been through this who can point me to a faster escalation path (or a specific team/contact), I’d be hugely grateful. submitted by /u/Zealousideal_End_366 to r/aws [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditgooglecloudimportance 0.52View on Reddit

For over two decades, agency hosting economics were beautifully predictable. You bought a reseller web server or dedicated cPanel account for $50 a month, crammed 30 client WordPress sites onto it, and charged each client a flat $25 monthly maintenance fee. Your margins were clear, your server bills were static, and billing surprises were virtually non-existent. Read the comple te article here > Serverless Bill Shock: Track Vercel & Supabase Client Costs | InstaRenewal Then came the modern web stack. Driven by the demand for lightning-fast digital experiences, agencies aggressively migrated to decoupled architectures: Next.js, Nuxt, Vercel, Supabase, Cloudflare Workers, and serverless databases like Neon. While the performance gains of this modern paradigm are undeniable, it introduced a chaotic operational reality: micro-subscription fragmentation and variable utility billing. submitted by /u/JadeLuxe to r/googlecloud [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.52View on Reddit

https://i.imgur.com/CPpTubs.png Well I guess it's a good day to test our backup datacenter. AC went out last night, at 3AM equipment started alerting rising temperatures. 5AM systems started shutting off so we moved to our backup and shut down everything else. submitted by /u/Pryach to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.52View on Reddit

From "RFC 1925: The Twelve Networking Truths" ( https://datatracker.ietf.org/doc/rfc1925/ ), published in 1996: «(6) It is easier to move a problem around, for example by moving it to a different part of the overall network architecture, than it is to solve it. (6a) Corollary: It is always possible to add another level of indirection.» Written for networking. Still uncomfortably accurate for multi-agent architecture. submitted by /u/MissionFinOps to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.52View on Reddit

Well it happened

by 1337DSSICTPDX

We had a critical outage due to a failing device that required me to go on-site and troubleshoot for six hours. We have zero documentation, so our standard process is to reach out to a senior team member for assistance. Their first suggestion was to reboot a switch, and then reboot the upstream switch. When that didn’t work, they had me repeat it twice. The cable management is so poor that you cannot read any indicators on the device, and the senior I was working with doesn’t know how to access or use that switch’s console. The only member with console knowledge wasn't available until noon. When they finally came online and I caught them up to speed, they considered the down devices to be upstream. My mind broke at that moment, and I’ve been spiraling ever since. We have no network topology maps, and having a department head flip-flop basic networking terminology is mind-blowing. Essentially, we had an outage at three locations, one of which had their production and sales affected. This whole thing could have been identified at home and quickly resolved by switching ports if we had proper documentation, which makes my heart sink. Am I overreacting? What would you all do in this Oh and to add insult to injury me asking about doing pir was laughed at. This is insaine. Fuck my life. submitted by /u/1337DSSICTPDX to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.52View on Reddit

When comparing OpenAI for a client, they expressed a strong preference for GPT models, so I evaluated directly using OpenAI's API, OpenRouter, and Azure. The company is already on Azure, so they said they prefer to go down that route, as it'd just be one single unified bill. Thing is though, when I compare OpenRouter pricing for Luna to Azure's pricing for the same model , the difference is 10x, 5x if you compare it to without the discount or OpenAI direct pricing What gives? Why is Microsoft charging 5x-10x the price of the same model elsewhere, am I getting something unbelievable I won't get anywhere else for that money, or is this just corporate tax? submitted by /u/aevitas to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X