CommPulse

CommPulse

1160 parked Settings

The cross-site community pulse: gold-layer posts + comment threads read live from the Communication Hub, ranked by importance. Turn a post into Discord / LinkedIn / X.

redditsysadminimportance 0.52View on Reddit

I inherited the building along with the server room. Here's how I'd pick an asset management system — 7 options, honestly compared Disclosure up front: I build one of the seven tools below (UniAsset, #3). I've written this so it's useful if you ignore that one entirely, and I've said where it loses. If that's not your thing, the other six are covered properly and you can skip the section. I want to describe a specific situation, because I don't think it's rare and I don't think it's well served. You're the IT person at a mid-sized organisation — a school, a charity, a hotel group, a manufacturer, a leisure centre. You started out responsible for laptops, switches and the Microsoft tenant. Then, over about three years, you also became responsible for: The AV kit that gets signed out and doesn't come back The vehicle fleet, because the keys live in your office The PAT testing records, because you own a label printer The fire extinguisher servicing dates, because the last person who tracked them left Whatever Finance means when they say "the asset register" None of this was a decision. It accumulated. And you're now managing it across three spreadsheets, a shared mailbox, and a calendar reminder you set in 2023 that you no longer trust. If that's you, the standard advice — "just use Snipe-IT" — is right about a third of the time and actively wrong the rest. What follows is the framework I'd use, then seven tools, then blunt verdicts. On ordering: the seven below are ordered by how far they sit from pure IT asset tracking toward full maintenance management . It is not a ranking. #1 is the best choice for a lot of people reading this and #7 is the best choice for others. Part 1: The question before the tool Almost every bad asset management purchase I've seen came from skipping this. People compare feature lists before they've decided which of four different problems they actually have. Problem A — "I don't know what we own." You need discovery and a register. Nothing else matters yet. Problem B — "I know what we own, but I don't know where it is or who has it." You need custody: check-in, check-out, assignment, audit trails. Problem C — "I know what we own, but it isn't being maintained." You need maintenance execution: schedules, work orders, someone getting a job on their phone. Problem D — "I know what we own, but I can't prove anything about it." Finance wants a depreciation schedule. Insurance wants serial numbers. An inspector wants current certificates. This is a compliance and evidence problem, not a tracking one. Most organisations have two of these. Very few have all four. The tools below are each genuinely good at one or two and mediocre at the rest — which is fine, as long as you match them to your actual problem rather than to a feature grid. Part 2: Five axes that actually decide it Everything else is noise. These are the five I'd score on. 1. Deployment. Do you have a VM and someone who will patch it in eighteen months' time? Be honest — "we can self-host" is often a statement about capability rather than about whether it will actually happen. Self-hosted is free until it's your weekend. 2. Discovery vs. manual entry. Does the tool find things on the network, or does it wait for you to type them in? This is a hard split. Network-discovered assets are IT assets. Everything else — the fridge, the minibus, the extinguisher, the treadmill — gets entered by a human, forever. If most of your estate isn't on the network, discovery is irrelevant to you and you should stop weighting it. 3. Maintenance depth. Recording that a service happened is a different product from scheduling recurring work, and both are different from dispatching a crew of thirty across a shift roster. Buy for the depth you have staff for, not the depth you aspire to. 4. Financial and compliance output. Can you export something Finance will accept without rebuilding it? Can you produce, on demand, every certificate expiring in the next sixty days? This is where most cheap tools stop and most people don't notice until an audit. 5. Who actually operates it daily. A dedicated maintenance administrator will happily learn a dense tool. An IT manager doing this alongside four other jobs will not, and neither will a caretaker on a phone in a plant room. Adoption failure is the number one cause of death for these systems and it's almost always an interface problem, not a feature problem. Part 3: The seven 1. Snipe-IT — the default, and often the right one Open source, self-hostable, mature, and genuinely good at what it does: an IT asset register with check-out to users, licence and accessory tracking, consumables, custom fields, and a decent API. There's a paid hosted option if you don't want to run it. Where it wins: IT hardware, one location or several, a person who's comfortable with PHP and MySQL, and a budget of zero. For "300 laptops, 40 monitors, who has what," it is hard to beat and you should probably just deploy it. Where it stops: it's an asset register, not a maintenance system. There's no meaningful preventive scheduling, no work order lifecycle, no inspection checklists. Financial reporting is thin — it holds purchase cost and depreciation fields, but it isn't producing a fixed asset register your auditor will sign off. And you own the upgrade path, the backups, and the security patching. Worth knowing: GLPI sits in adjacent territory — also open source, also free, with a helpdesk bolted on. If you want an asset register and ticketing without licence spend, look at it before you settle on Snipe-IT. It's heavier and the UI is an acquired taste. 2. Lansweeper — when you don't trust your own inventory Different category. Lansweeper's core competence is finding things: credential-based network scanning, agent and agentless options, software inventory, OS and patch state. It answers "what is actually on this network right now," which is a question no manual tool can answer. Where it wins: you've inherited an estate nobody documented, or you need software inventory and licence position for a true-up. Discovery is a genuinely hard technical problem and Lansweeper solved it properly. Where it stops: it's a discovery and inventory platform, not an operations platform. Custody workflows and maintenance scheduling are not what it's for, and anything not on the network — which is most of a building — is invisible to it. Plenty of shops run Lansweeper for discovery and something else for everything downstream, and that's a reasonable architecture rather than a failure. 3. UniAsset — assets and their paperwork in one record This is mine. Read it sceptically. The gap I built it for is the one in Problem D above: organisations where the assets and the documents about the assets are the same job. The gas safety certificate, the calibration record, the insurance schedule, the warranty, the vehicle MOT. Almost nothing in the affordable tier treats document expiry as a first-class monitored condition rather than as file storage. What it does: asset register with hierarchical locations and categories; documents and images attached to assets with expiry dates that generate warnings ahead of the date and again once passed; preventive maintenance schedules; work orders with checklists, materials and photos; check-out custody separate from durable assignment; depreciation across six methods with a Fixed Asset Register export; total cost of ownership per asset; an append-only event log on every asset; and incident records that snapshot the asset's state at the moment of the event and lock it, which is what makes them usable in an insurance claim. Five roles, multi-site, QR codes generated per asset, installable as a PWA so field staff don't need an app store. Where it genuinely wins: the mixed estate. IT hardware sitting in the same system as building plant, vehicles and equipment, where Finance wants a depreciation schedule and someone wants proof that the certificates are current — operated by one or two people who have other jobs. Where it loses, plainly: No network discovery. No agent. Assets are entered manually, by CSV import, or through the API. If Problem A is your problem, use Lansweeper. No software licence management. It does not track licence entitlements or do a true-up. That's a real gap versus Snipe-IT for a pure IT use case. No self-hosting. Multi-tenant SaaS only, and that's a deliberate decision rather than a roadmap item. If self-hosting is a requirement, this is disqualified — stop here. Not an ITSM. No ticket queues, no change management, no service catalogue. The SLA features apply to maintenance work orders, not to service requests. No dispatch optimisation or shift scheduling. If you're routing a crew of thirty, a real CMMS beats it. IoT ingestion exists but nothing reacts to it. Signals are stored and displayed on the asset timeline; there's no rules engine turning a temperature threshold into a work order yet. If you want signal-driven automation today, pair it with a dedicated IoT platform or pick something else. Pricing shape: free tier with unlimited assets — asset count is deliberately not the meter — but capped at 5 users, with the register, documents and expiry alerts included. Maintenance scheduling, work orders, custom fields and the depreciation/FAR layer start at the first paid tier. API and webhooks are a tier above that. There's a nonprofit plan for qualifying organisations. So: if you only have IT hardware, #1 is better and free. If you don't know what you have, #2 is better. This is for the case where the certificate expiring matters as much as the laptop going missing. 4. EZOfficeInventory — when things move constantly SaaS, built around the check-in/check-out cycle. Barcode and QR scanning, reservations and bookings, custody history, a maintenance module, decent mobile apps. Where it wins: shared equipment that circulates all day. AV gear, tools, loaner kit, camera equipment, test instruments. If your dominant pain is "who has the thing and when is it coming back," this is purpose-built for exactly that and it shows. Where it stops: the maintenance side is present but light compared to a real CMMS, and the financial reporting is serviceable rather than audit-grade. Per-user pricing means it gets expensive if lots of people need to touch it — worth modelling before you commit. 5. Asset Panda — when every asset needs its own shape Extremely configurable. You define the fields, the layouts, the workflows, the relationships between record types, and you get a solid mobile app on top. Effectively a structured database that arrives knowing it's about assets. Where it wins: non-IT estates with genuinely heterogeneous records — where a vehicle, a defibrillator and a piece of stage lighting all need entirely different field sets — and where field staff are on phones. Where it stops: configurability is a cost as well as a feature. Someone has to design and maintain the configuration, and there's a real implementation effort before it's useful. Pricing is quote-based, which makes quick comparison hard. If you want opinionated defaults rather than a blank canvas, you'll find it heavy going. 6. MaintainX — maintenance, from a phone Mobile-first CMMS. Work orders, procedures and digital checklists, asset records, parts, messaging built into the work order. Designed for people in a plant room or on a floor, not at a desk. Where it wins: you have maintenance staff and the problem is that work isn't being logged or completed. Adoption is the thing it's actually optimised for, and it's noticeably better at that than tools that grew from a desktop-first design. Where it stops: it's a maintenance system with asset records attached, not an asset operations system. Depreciation, fixed asset registers and document expiry compliance are not its centre of gravity. If Finance is the requester, this isn't the answer. 7. Limble — a real maintenance department Full CMMS. PM scheduling, work request portals, asset hierarchies, parts and inventory, meter-based triggers, and reporting deep enough to run a maintenance function on. Widely liked in this space, and deservedly. Where it wins: you have a maintenance supervisor, a technician team, a backlog, and someone who reports on downtime and PM compliance. This is what it's built for and it's very good at it. Where it stops: it is more system than a one-person IT-and-facilities operation will ever use, and pricing reflects that. Buying a CMMS because you have twelve PM schedules is buying a lorry to move a sofa. Part 4: The comparison Deployment Network discovery Free tier PM scheduling Work orders Document expiry alerts Depreciation / FAR Check-out custody Software licences Snipe-IT Self-host or cloud No Yes (self-host) No No No Basic fields Yes Lansweeper Self-host or cloud Yes — best in class Limited No No No No No UniAsset SaaS only No Yes (5 users) Yes Yes Yes Yes Yes EZOfficeInventory SaaS No Trial Light Light Partial Basic Yes — best in class Asset Panda SaaS No Trial Configurable Configurable Configurable Configurable Yes MaintainX SaaS No Yes Yes Yes — best in class No No Light Limble SaaS No Trial Yes — best in class Yes No No Light Two columns are worth staring at. Discovery has exactly one real answer, and if you need it you need it. Document expiry alerting is the column most people don't think to ask about and then discover they needed — a folder of PDFs is storage; a thing that tells you sixty days out is compliance. Part 5: Situation → verdict Blunt, no hedging. "Only IT hardware. No budget. We have someone who can run a VM." → Snipe-IT. Stop reading, deploy it this week. "I genuinely don't know what's on the network." → Lansweeper , and pair it with something else downstream. Discovery is its own problem. "Asset register plus a helpdesk, no licence spend." → GLPI. "Assets and their certificates both matter. Finance wants a depreciation schedule. One or two of us run all of it alongside another job." → UniAsset. This is the case I built it for, and if it isn't your case one of the others is better. "Gear circulates constantly and doesn't come back." → EZOfficeInventory. "Every asset type needs a completely different field set, and staff work from phones." → Asset Panda , if you have someone to own the configuration. "We have maintenance staff and the work isn't getting logged." → MaintainX. "Real maintenance department, real backlog, real PM compliance reporting." → Limble. "Regulated maintenance regime, deep ERP integration, procurement in scope." → None of these. Maximo or SAP PM , and budget two quarters and a consultant. "Under 50 assets, one site, nothing expires, nobody audits us." → Honestly, a spreadsheet, and put the money somewhere it matters. Everything above is overhead until the estate or the accountability grows. Part 6: Five questions I'd ask any vendor These separate serious products from demos, and none of them appear on a feature grid. Is the pricing metered on asset count? If yes, understand what happens when you import everything on day one — it's the single most common reason a rollout stalls halfway through data entry. Can history be edited? Ask specifically whether the change log is append-only or whether an admin can quietly alter a past record. If you'll ever use this for an insurance claim or an audit, an editable history is worth very little. What happens to my data if I downgrade or cancel? Are custom fields deleted? API keys revoked? Get it in writing. Can Finance use the output directly? Ask for a sample depreciation schedule or fixed asset register export and show it to your accountant before you buy, not after. What's the import path, realistically? Ask what percentage of implementations complete data import. The answer will be evasive, which tells you something. Part 7: What actually kills these projects Not the software. Three things, in order: Data entry. Everyone underestimates it. 500 assets at two minutes each is seventeen hours of someone's life, and it always comes out of evenings. Whatever you pick, get the CSV import working properly before you enter a single record by hand, and accept a partial register that's accurate over a complete one that's guessed. Taxonomy. If you don't agree the category structure before you import, you'll have "Laptop", "Laptops" and "Notebook" within a month, and every report will be wrong in a way nobody notices. Decide the hierarchy first, and make sure the tool resolves category names against a list rather than creating them from whatever gets typed. The single point of failure. These systems get built by one enthusiastic person and abandoned when they leave. If only one person knows how the thing works, you've replaced a spreadsheet with a more expensive spreadsheet. Make sure at least two people can operate it and that the data can come back out. Happy to answer questions on any of the seven, including the awkward ones about mine. And if you've deployed one of these and it didn't go the way this post suggests, say so — that's more useful to whoever reads this next than another feature list. submitted by /u/ishrargo to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditgooglecloudimportance 0.52View on Reddit

Same unrestricted Gemini key hole as the $82k stories, later in 2026, smaller bill. Google's own CLI still minted API keys with no restrictions. Those keys could call billed Gemini. Truffle Security reported that class of defect on 21 Nov 2025. Google first treated it as intended behaviour, then as a bug. They did not start rejecting unrestricted keys for Gemini until 19 June 2026. On this account: 53,197 Gemini Generative Language requests between 24 May and 6 July 2026, peaking at 12,693 on 5 June. The spike died the day they flipped that block. Not a stolen Google password. An unrestricted key their tool created by default. How they responded: they cited July charges of $4,446.45, applied a 75% "maximum goodwill" credit of $3,257.16, and still demanded about $1,190 to $1,300 as a Cloud balance after a $4,000 card chargeback. Closeout blamed an unauthorized party "linking external projects." Live billing IAM at the time did not show that. Budget alerts are not a cap. They sat on the Truffle report for months, then kept a quarter of a bill that only exists because of the hole they eventually closed. No names, no project ids, no keys. submitted by /u/unrestrictedkey to r/googlecloud [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.52View on Reddit

Follow up to my last post, this time I did it on purpose. Put together a "typical prod mistakes" stack: previous gen instances, gp2 everywhere, an orphaned EBS volume, a full size database duplicated into staging and left running 24/7. F, 0/100, about $2,694/mo, with $1,251/mo flagged in savings across 15 recommendations. Image is exported directly from the tool's own Export PNG button, wanted it to look exactly like what you'd see running it yourself, not a mockup. submitted by /u/Independent-Ease-609 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditdevopsimportance 0.52View on Reddit

Something that keeps coming up during our incident response is just how much time we lose jumping between the ide and whatever tool holds the relevant logs, traces, or metrics. Typical flow: you are in the ide looking at a failing code path, you hit unexpected behavior and the next 20 minutes is alt-tabbing between your editor, log search, a distributed tracing ui, metrics dashboards, feature flag console and deploy history. You copy a trace id from logs over to the tracing tool then you copy a user id back into a sql query then you try to map all of that back to the exact function and commit you are staring at in the ide. We've got what most people would call a modern observability stack: distributed tracing, structured logs, dashboards, decent tagging and reasonably instrumented services. the problem isn't that the telemetry doesn't exist, it's that none of it really lives where developers spend their time writing and reviewing code. During incidents, people end up doing their own ad‑hoc integration work: copy from log search, paste into the ide, grep locally, jump back to the metrics dashboard, repeat. The pain points i keep seeing during production debugging are pretty consistent. there's no single place that shows this line of code, these commits, these deploys and these recent errors and traces in one view. Most observability tools are optimized for operators staring at dashboards, not developers trying to understand how a specific code path behaves in production. even when telemetry is tagged correctly, you still have to remember which query or dashboard to open and how to line it up with what you're debugging in the ide and during a live incident, that context‑switching overhead turns directly into mttr and oncall fatigue. What's interesting is that we keep buying more observability tooling but the core developer workflow is still: ide here, production reality over there and your brain plus clipboard as the glue connecting the two. How have you cut down on context switching between the ide and your logs, traces and metrics during debugging and incident response, whether that's pulling production context directly into the ide, pushing more code context into your observability tools or standardizing on a single pane for incident work? submitted by /u/Potential_Force_4136 to r/devops [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.52View on Reddit

Half your reliability spend is on systems nobody would notice going down. I audited a shop last year with a $180K a month multi-region setup on an internal reporting tool used by roughly 12 people on the finance team. Four nines of availability. Automated cross-region failover. Full DR runbooks. When I asked the CFO what happens if it goes down for 4 hours on a Tuesday, he shrugged: "we get the numbers Wednesday." That's not a cost problem, it's a decision problem. The infra was rightsized. Tags were clean. Utilization looked healthy on every dashboard. The waste was invisible because well-utilized well-tagged well-sized infra IS the metric everyone watches. Cost dashboards surface the wrong axis. They show you spend, utilization, waste-per-service. They don't answer "how much of this uptime does anyone actually need?" The tradeoff between availability and cost only exists as a decision, and once it calcifies into an architecture nobody revisits, the over-reliability spend just compounds. Common pattern: a "drop this workload to one AZ, save $60K a year" recommendation lands in a ticket. The engineer who owns it sees the cost side, can't see the reliability side of the tradeoff, and marks it "no, we need HA" without ever asking the business side what HA is actually worth. Anyone else seeing over-reliability as a category of spend you can't touch without reopening the original architecture decision? submitted by /u/matiascoca to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.52View on Reddit

Hey everyone, Wanted to share a fix for a frustrating compute allocation issue we ran into recently with Azure Batch and Azure Data Factory (ADF) that might save some of you a few thousand dollars on your monthly Azure bill. The Problem We were running hundreds of Python/Pandas data pipelines triggered via ADF—a mix of scheduled batch runs and sequential dependent chains. To handle the load, we scaled out an Azure Batch pool with 4x Standard_E4ds_v5 VMs (16 total vCPUs). On paper, everything looked sized correctly. In reality, our Azure Metrics showed CPU utilization hovering at ~25% , while scheduled tasks queued up. 12 out of 16 vCPUs were sitting completely idle, even though we were paying full price for all of them. Why It Happens The issue comes down to a structural mismatch between Python, Pandas, and Azure Batch defaults: Python GIL & Pandas: Standard Python uses the Global Interpreter Lock (GIL), and Pandas operations are natively single-threaded. One Pandas script can only ever utilize 1 CPU core . Azure Batch Gatekeeper Default: Unlike a local laptop OS (which aggressively context-switches and interleaves processes), Azure Batch defaults to Task slots per node = 1 . The Result: Azure Batch locks an entire 4-vCPU VM for one Python script. That script uses 1 core (25%), while Azure Batch walls off the remaining 3 cores from accepting any queued ADF tasks. The Fix Depending on your data scale, there are two ways to fix this: High file volume / smaller files (Scale ACROSS tasks): Recreate your Batch pool with Task slots per node = 4 (matching your VM core count). Set Node Fill Type to Pack . Result: Each 4-core VM runs 4 independent Python tasks simultaneously. Zero code changes required, and CPU utilization hits 100%. Single massive datasets (Scale WITHIN the task): Keep Task slots per node = 1 . Swap single-threaded Pandas for Polars ( import polars as pl ) or Modin . Polars uses a multi-threaded Rust engine that automatically parallelizes across all 4 vCPUs for a single script. Cost & Resilience Bonus We also paired this with a hybrid pool strategy ( 20% Dedicated baseline for SLA protection + 80% Low-Priority/Spot for cost savings) alongside a retry policy on our ADF custom activities to handle node evictions cleanly. I wrote up a detailed post explaining the underlying mechanics, decision matrices for file sizing, and pool configurations here: 🔗 Stop Paying for Idle Compute: How We Fixed the Python and Pandas Bottleneck in Azure Batch Curious if anyone else has run into this default task slot trap, or if you're taking a different approach (like running PySpark/Databricks instead) for these types of workloads submitted by /u/straightfromthegut to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.52View on Reddit

Woke up late today; noticed I had no emails since 5 AM (PST); went to their website and read this severe notice to move off MXGD asap. MXGuarddog.com Maintenance Mode MXGuarddog Large Scale Outage Our primary data center, located in Phoenix, AZ, has suffered a total failure of its cooling systems. The facility is 200,000 square feet, and it is hot—really hot. We started to receive warnings from our monitoring systems that temperatures in the data center were critical. Within minutes, we lost all connectivity to the facility. We do not own this data center; we colocate our equipment in the facility. The data center has not provided any ETA for when they will be back online. They have only advised that their technicians are working to restore the cooling systems. As a result, we do not know when we will be able to bring our systems back online. At this time, we strongly recommend that any MX Guarddog customers who can do so update their MX records to point directly to their own mail servers in order to bypass this widespread outage. We will provide additional information as soon as we receive it. Here are updates we are receivig from the data center: All hands are on deck working with our datacenter provider in our Phoenix node. Currently, the datacenter provider has on-site vendors working to resolve the situation and restore adequate cooling capacity. Some capacity appears to have been restored. Their datacenter technicians have established mitigation efforts to reduce rising temperatures. At present, we have no eta on full resolution. submitted by /u/TechnicianOnline to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.52View on Reddit

Hey all, Every time I read about downtime for an airline, inevitably the tech discussion in the comments boils down to how some of these systems are 50+ years old. I get it, don't fix what isn't broken, and minutes of downtime can translate to $$$ of losses. But I'm wondering if there's any of you that work in the industry who have encountered an actually good system that you like and works really well? Or are these systems just things you build up muscle memory for and ultimately just get used to? I wanted to go hunting for examples of this kind of software to learn more about it but I literally don't even know what the search terms would be. submitted by /u/CtrlAltDelve to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.52View on Reddit

Honest question, and a disclosure to start with: I am building a tool in this area, so I am biased. I just want to know whether the problem I think exists actually exists for you. Here is the gap as I see it. Most organisations I have worked in my careear as a consultant had good coverage of the two extremes: * Monitoring (Zabbix, PRTG, Checkmk, etc.): tells you that host X is down, right now. * Inventory (Lansweeper, Snipe-IT, Intune, a scanner of some kind): tells you what is installed on host X. What nobody seemed to have was the middle layer: host X runs application Y, application Y is the backend for service Z, service Z is used by finance and by two other applications through some kind of integration, and the person who owns Z is not available right now ( holiday etc). That knowledge lived in three heads and a wiki page no longer maintained. So when something went down, the process was: alert fires, someone asks "what is that box for?", someone else asks "who do we need to tell?", and the answer was assembled from memory, wikies, excel sheets and so on. The CMDB was supposed to fix this. In my experience it never did, because keeping the relationships current was manual work nobody really did, so it drifted, nobody really knew. What I would like to know from people who run this stuff: When a server, a database or an integration fails, how do you today find out which applications and which people are affected? Is it tooling, documentation, or common knowledge? Does anyone actually keep application-to-server and application-to-application dependencies up to date? If yes, how, and what made it stick? Has NIS2 (in EU) or an audit, or an insurer forced you to produce this map? What did you hand over, and how painful was it? If a tool did this, what would it need to do to be worth the effort, rather than being another CMDB that is kind of forgotton? My current assumption is: it has to discover as much as possible itself (what runs where, what talks to what), and it has to be honest about what it does not know instead of pretending the infrasrtruture map is complete. And if in your organisation , the answer is "we do monitoring and keep the other parts in spreadsheet is fine, this is a solved problem", I would genuinely like to hear that too. That is a cheaper lesson to learn now than later, i dont want to finalize a product no one needs. submitted by /u/michaelnielsen to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.51View on Reddit

Curious how teams are handling this. Between OpenAI/Anthropic API bills, GPU instances on RunPod or EC2, and vector DB costs, it seems like most places have one big number and no idea which feature or model is driving it. Do you know your cost per request, or per user, for anything AI-powered? Is anyone tracking token spend by feature, or is it all one line item? If you self-host, do you know your actual GPU utilisation, or is it "the box is up"? Has anyone gone through and actually cut this — what worked? submitted by /u/kondamuri [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.51View on Reddit

Genuinely curious how other teams handle this, because at every place I've seen it works differently and mostly badly. The bill comes in, someone says it's too high, and then... what? Is there one person whose job it is to dig in? Does it fall on whoever is least busy that week? Does it just get ignored until finance escalates? Specifically wondering: - Is anyone on your team formally responsible for cloud spend, or does it float? - When the bill jumps 20% month over month, how do you find out why? Dashboards, a tool, or someone manually clicking through Cost Explorer for two hours? - Do you review it on a schedule, or only when someone panics? - Has anyone actually done a proper line-by-line audit, and did it stick or did spend creep back up in six months? Mostly asking because I keep seeing the same pattern — everyone knows the bill is too high, nobody owns it, so nothing happens. Curious whether that's universal or whether some teams have solved it. submitted by /u/kondamuri [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.50View on Reddit

I was building a FinOps view for an AI support workflow and kept staring at the word productivity. Request count, spend, and latency all fit into tidy charts. None of them told me whether a support case stayed closed or came back two days later. That is awkward when the dashboard is supposed to tell me whether the workflow is helping. Then I ran into two recent studies that seemed to disagree. Firm Data on AI surveyed nearly 6,000 executives, and 89 percent reported no productivity impact over the previous three years. AI, productivity, and the workforce used a sample of nearly 750 executives and found positive but uneven gains. The samples and questions differ, so it is not a clean contradiction. I came away thinking that the answer depends heavily on what question you asked in the first place. For this support workflow, I need data from both sides. The provider bill tells me the inference cost. The help desk has handling time and reopened cases. Looking at either one alone is like checking the grocery receipt without asking whether dinner was edible. A cheap draft can still create a lot of cleanup for the person reviewing it. Suppose a support rep spends five minutes prompting the system and gets a draft, then spends another 40 minutes checking and rewriting the answer. That is 45 minutes of human time plus the inference cost. If the dashboard records only the first five minutes, it gives the model credit for work the rep had to redo. A reopened case would make that picture worse. For a first pass, I can take 25 AI-assisted cases and 25 unassisted cases from the same support queue and week, then match them by issue type and the support rep's experience as closely as I can. ZenMux can give me cost and latency for the assisted requests. The help desk supplies the less glamorous half of the story: handling time and reopens. Fifty cases will not settle a company-wide argument, but they are enough to see whether the current dashboard is flattering us. Even then, I will not have proof that AI caused whatever difference I find. I will know whether repair time and reopened cases wipe out the savings I thought I had. If the workflow only looks productive when I ignore the human cleanup, I need to fix the measurement before I expand it. submitted by /u/quietkida7 [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.46View on Reddit

I keep seeing "cost per customer" thrown around like it's a simple metric, but once you factor in shared infra like RDS or a shared EKS cluster, the attribution gets messy fast. Anyone got a practical framework for splitting shared resource cost across tenants without it turning into a spreadsheet nightmare? Would love to hear how teams handle this in practice, not just in theory. submitted by /u/yourcloudguy [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X