Maintaining Apache Iceberg Tables: Compaction, Snapshots, Metadata and Orphan Files
by codingdecently
submitted by /u/codingdecently to r/FinOps [link] [comments]
The cross-site community pulse: gold-layer posts + comment threads read live from the Communication Hub, ranked by importance. Turn a post into Discord / LinkedIn / X.
by codingdecently
submitted by /u/codingdecently to r/FinOps [link] [comments]
by grishma_1503
when you've started rotating a credential and then discovered it was already revoked, or was never a real key in the first place — roughly how long had you burned before you worked that out? minutes, or did someone spend an hour convinced they were looking in the wrong project? submitted by /u/grishma_1503 to r/devops [link] [comments]
by Pope_Carl_LXIX
Company is thinking about getting onboard the AI train exploring options. What are some questions that should be answered throughout the review process and things to know going forward when implementation is complete? submitted by /u/Pope_Carl_LXIX to r/FinOps [link] [comments]
by Slight_Piccolo_3746
Recently we tried to cut the model price on an internal research agent and it barely moved the needle. We initially had the obvious theory that the expensive model caused the spend. When we then looked at cost per successful task instead of cost per model call we saw something different. A large chunk of the sessions called search repeatedly, pulled nearly identical results, and stuffed every tool output back into context. Some heavy sessions were legitimate. Complex research really did need multiple searches and a larger model. Others were just retry loops caused by weak stopping conditions, bad tool arguments, or the model missing that the previous result already answered the question. The cheaper model occasionally made that worse by needing more attempts. We used Braintrust to break down token counts, latency, cost and tool-call patterns by trace. That gave product, engineering and finance the same unit of analysis, a completed task rather than a single API request. We added loop detection and deduplicated tool outputs then routed simple requests to the smaller model and compared the change against the same eval set. What we saw was fewer repeated calls which drove more savings than the model swap and answer quality held. Latency improved too because the agent stopped arguing with the search endpoint five times. How are people measuring agent cost when retries and tool loops make per-call metrics basically meaningless? submitted by /u/Slight_Piccolo_3746 to r/googlecloud [link] [comments]
by VaporousMote
Small business. Not a huge budget, especially not given the economic situation. We run Meraki MX devices at a number of offices. We also have a number of creative users that work from home and need fast + resilient tunnels to navigate and transfer files quickly. Semi-recently Meraki began offering IKEv2/IPSEC. Works great, super fast. Problem is, they don't have MFA support for it yet. They seem to want you to use AnyConnect Premium, but the Meraki only supports TLS/DTLS tunnel type, which is substantially slower for transferring large files. They also don't support client certificate validation for IKEv2/IPsec, which would be another good option that isn't "anyone with a username/password who knows the termination IP/name can connect." Right now we authenticate by pointing the MXs at an NPS server in Azure, which is joined to an Entra DS domain. We want to avoid managing an AD domain but are heavily integrated into the MS ecosystem (Teams, Office, Win11 Business, Intune, etc). Is there a tool or service that we could point the MX's RADIUS server field at, that integrates with Entra/Entra DS that could perform MFA on its end before returning success and granting access, augmenting the basic username/password auth? submitted by /u/VaporousMote to r/sysadmin [link] [comments]
Two weeks ago I posted my concerns about the long term viability of using Gemini in my automated content moderation app. We covered the option of alternate models there so please constrain this discussion to what appears to be sneaky policy settings to extract more revenue from Gemini users. When 2.5-flash-lite is deprecated my costs will be 16x on 3.1-flash-lite but actually it will be more than that because on the 3 series models you can't disable thinking nor specify temperature, top_p and top_k. Of those I only know a little about temperature which I have set set to 0.1 for my purposes. From 3.1 onward the best I will be able to do is specify a "LOW" thinking level which will burn at least some additional tokens and might break my app because having the model act deterministically is essential for content moderation purposes. Sure I could add additional system instructions to try and compensate for this loss of control but up go my input costs. Granted you can still specify maxOutputTokens but if that value is too low to account for the mandatory "thinking" the call will fail is that right? So we're going from a situation where you can tightly control the cost of each call to the LLM to one where you're at the mercy of the model. As I mentioned above I think this policy is sneaky which would be entirely consistent Google's opaque cloud billing and costs in general. Alternatively these changes are just where the rubber meets the road? submitted by /u/dougception to r/googlecloud [link] [comments]
by Real-Patriot-1128
I was in the process of creating an Azure SQL Managed Instance for our CIS dept (I work at a College). The cost estimate came out to $1,400/month - which is too much. Any way around that? submitted by /u/Real-Patriot-1128 to r/AZURE [link] [comments]
by BB03440
Last update August 14, 2026 at 8:29am PDT Description At 8/14/2026 7:35 AM PT, the Core Identity team became aware of an issue with email deliverability affecting customers in commercial cells. During this time, users may be experiencing intermittent issues receiving email from Okta. Our team is actively investigating this issue and is working to mitigate it. 8:39am PDT: Okta Engineering has determined that certain third-party providers were failing to receive emails from Okta. Our monitoring shows delivery rates are recovering to normal conditions and will continue to monitor until full resolution. Our next update will be in 30 minutes or sooner if additional information becomes available. 8:23am PDT: Okta Engineering is investigating and has determined that the incident is currently impacting email delivery to a subset of customers. During this time, customers may experience email delivery deferrals. We’ll provide an update in 30 minutes, or sooner if additional information becomes available. https://status.okta.com/#incident/a9CWR0000002FZG2A2 submitted by /u/BB03440 to r/sysadmin [link] [comments]
by cloudquell123
[ Removed by Reddit on account of violating the content policy . ] submitted by /u/cloudquell123 to r/FinOps [link] [comments]
by ui-consulting
When I talk to analytics leaders, the #1 thing they complain about is data debt. They assume their pipelines are fundamentally broken and that the only fix is a six-figure, multi-month overhaul of their cloud stack. Usually, that’s completely wrong. The problem isn't that data debt exists—every scaling company accumulates it as a natural byproduct of growth. The real problem is that their data debt is completely hidden in the dark . When data debt lives in the dark: Executives sit in 9:30 AM daily syncs arguing over spreadsheet semantics and metric definitions instead of making strategic decisions. Stakeholders treat the analytics team like a fast-food drive-thru, shouting isolated data requests into the microphone without context. Analysts waste 80% of their bandwidth acting as "data detectives" trying to trace broken C-suite CSV exports. At U&I Consulting , we don't promise to magically erase data debt overnight. Instead, we illuminate it . We bring hidden operational pipeline deficiencies into the light using Data Debt Diagnostics and formal Data Debt Tickets . When you quantify data debt and map it visually, it transforms from an invisible bottleneck into a clear, business-justified roadmap for the executive team. By combining this with Agile Ledger Architecture (ALA) , Medallion Pipelines (Bronze/Silver/Gold) , and KPI Shields , you protect team bandwidth and accelerate executive Time to Insight (TTI) . Stop trying to pretend data debt doesn't exist. Illuminate it, quantify it, and build an architecture that lets your business scale past it. 📖 Detailed in my book, "WHERE ARE THE INSIGHTS? The Blueprint for Agile Ledger Architecture" 🌐 Advisory & Diagnostics: uiconsulting.com submitted by /u/ui-consulting to r/FinOps [link] [comments]
by simply_complexiyo
Hello So i will be quick i want to ask some questions which are -------------------------------------------------------------------------- 1- I am a complete begineer who knows nothing so how can i learn finops like what things or tools or softwares would i need to learn communications and other skills or are tools enough -------------------------------------------------------------------------- 2- Is a FinOps a smart move right now for somone who doesn't have any experience with cloud but is willing to learn also are finding jobs preferably easy and how much can you expect 3- Is a fully remote job possible -------------------------------------------------------------------------- 4-Would it be a good option right now as compared to other options -------------------------------------------------------------------------- i had other questions which i wanted to ask but remember so will ask them at a later time submitted by /u/simply_complexiyo to r/FinOps [link] [comments]
Anybody here who feels AI SRE is a gimmick? I mean, what are we doing with this, honestly? Not gonna name names, but our observability platform ships one, and it takes 10–12 minutes to run an incident analysis and comes back with "hey, these are 2 things that are broken, and these are 5 probable causes." In the last three incidents, the causes it flagged were real issues acting up as bugs, but not related to the current incident at all. So I'm getting a faster RCA, just a wrong one. I raise a ticket and get misleading suggestions. Just to get an unbiased opinion, I tried out a few other names on the market: similar results. Happy to hear if anyone has had a different experience. submitted by /u/Less_Ad8195 to r/devops [link] [comments]
We're seeing issues with Ubuntu clients (22.04/24.04/26.04) failing to check in with Intune since thursday. Intune logs on the clients say "Failed to Check in with Intune. Failed to update device inventory. <some MS URL> Status Code 500". Some other logs also say 400 or 404. Latest version of intune-portal and microsoft-identity-broker installed. Downgrade did not help. Trying to enroll a new device will either: result in the device only landing in Entra, but not in Intune result in the device appearing in both Entra and Intune, but without any inventory data Entra shows "Ubuntu+26.04+LTS" as OS version, while Intune shows either nothing (blank) or "0.0.0.0", resulting in the device not picking up our compliance policy -> device will be not compliant -> user cant login to M365 apps. dsreg on the client shows that it correctly registered to Entra. Trying to trigger a sync or log in locally in the Intune portal app will immediately result in an error message ("Device could not be registered. There was an error, please try again"). Existing clients that are already enrolled also fail checkin and logon on the Intune portal app. But since they're still compliant users can still log in to M365 apps. There seems to be an issue going on at Microsoft's side. There's an open issue in the admin center (OP1459987), which sounds like it might be related giving the intune-portal app fails to update its device inventory. "Some admins may see outdated device inventory and update status information in the Microsoft 365 Apps admin center" submitted by /u/EpicSimon to r/sysadmin [link] [comments]
by FIRE0118999881999119
We have a currently CSP providing us licensing with a 10% discount over Microsoft's direct pricing, and thats fine by us. This relationship started out great with having a dedicated person we could ask licensing questions to that was fairly knowledgeable, so that was a nice perk when navigating the licenses available became a bit too much. Unfortunately we've began running into billing issues. These issues are taking far too long to work out. Our contacts have also become fairly unresponsive to our emails taking weeks to get back to us at times. This all prompts us to start looking around at options. Does anyone have any companies they currently have a good experience with? I don't really want to cold call, find a CSP that promises everything, and then puts us right back where we're at right now. We're just looking for licensing with a discount over going direct via Microsoft, this should be purely transactional so we need someone that makes it that easy. We are in the ~150 user range since I know that can make a large difference. Located west coast US. Thanks! submitted by /u/FIRE0118999881999119 to r/sysadmin [link] [comments]
by WancloudsInc
submitted by /u/WancloudsInc to r/FinOps [link] [comments]
by doctorevil30564
I'm sitting here scratching my head trying to figure this one out, I've went through our group policy settings multiple times and I can't find anything I might have configured that would cause this problem. new out of the box lenovo thinkpad laptop with 25H2 preload on it from factory. works great until it is joined to the domain, as soon as it joins the domain and rebooted, subsequent logins for any domain account or local account on the machine are having issues with not being able to run any apps other than edge or file explorer and on login it's giving an error stating that your system administrator has blocked the program. when I check the logs I'm seeing DistributedCOM errors with event id 10001 for Microsoft.AADBrokerPlugin, that appear to be related to windows security core background get token task classid webaccountprovider being unavailable. I'm not sure what the heck is going on here, but I need to get it fixed before it spreads to any of our existing windows 11 machines if it was caused by a malfunctioning windows update or something else. I'm about to blow the machine away and just load a clean install of 24h2 on it, but if anyone knows how to go about fixing this I'd like to try that first before I give up. this one is a replacement laptop for an employee and he can manage for a few days with his current laptop. my google-fu skiils haven't came up with anything that has worked thus far on it. I've reset it and it runs fine again until it is joined to our domain. I did notice when I ran systeminfo from a command prompt it is reporting that App Control for Business policy is enabled and app control for business user mode policy is set to audit. Looking on another machine that is still working fine, those two settings aren't activated. I tried local group policy to turn device guard off to see if that made a difference, but it didn't. submitted by /u/doctorevil30564 to r/sysadmin [link] [comments]
by BLWHpurple
I currently handle SaaS & AI procurement for client orgs and want to add FinOps to my skillset. I'm non-technical, was in sales before procurement, and have a decent grasp on financial concepts having passed CFA level I (from a previous career path I thought I would go down but didn't). I have minimal cloud knowledge, should I start with AWS CCP to understand cloud fundamentals and so better understand the scenarios FOCP questions pose later, or start with the (as I understand it) less technical FOCP for a gentler learning curve? submitted by /u/BLWHpurple to r/FinOps [link] [comments]
https://preview.redd.it/4ceym8fdealh1.png?width=1672&format=png&auto=webp&s=839cc0b352c7b923cacd02b4386bc834f48dc47d A lot of compliance work isn't spent fixing security problems. It's spent finding the evidence . Screenshots. Spreadsheets. Configuration exports. Change logs. Emails between IT, security, risk, and compliance teams. Then the cycle starts again before the next audit. For Saudi enterprises working across frameworks such as NCA ECC, SAMA, PDPL, ISO 27001, PCI DSS, SOC 2, and others, the operational burden can become enormous. We broke down a five-step approach to making compliance more continuous: Map infrastructure to relevant frameworks continuously Generate evidence automatically Detect configuration drift early Unify visibility across multi-vendor environments Use AI for policy and gap analysis—not just reporting The interesting shift is from: "We need to prepare for the audit." to: "We should already be able to prove our compliance posture." That's where AI-driven continuous monitoring could have a significant impact. Full breakdown: How to Cut Compliance Audit Effort by 90%: A Practical Framework for Saudi Enterprises For anyone working in security or compliance: what currently consumes the most time during your audit preparation? #Compliance #Cybersecurity #NCAECC #SaudiArabia #AgenticAI #GRC submitted by /u/WancloudsInc to r/FinOps [link] [comments]
by SmartWeb2711
We run an internal self-service platform API. Individual humans wanting CLI/scripted access this is where I'm stuck . Currently we mint a static token, show it once in the UI, and the user keeps it. Two things bother me: Authority is frozen at creation. We store the list of accounts the token may touch. If the user later loses their admin role on one of those accounts, the token keeps working. The credential outlives the entitlement. Distribution is copy-paste. It ends up in .env files, shell history, occasionally a chat message. If you've done per-request authorization lookups, what did it cost you in latency and directory load? How long do you cache, and how do you handle the lookup failing fail open or fail closed? For humans needing programmatic access, has anyone made short-lived tokens work by exchanging an existing SSO session? Feels like the "right" answer but I haven't seen it described much outside cloud provider SDKs. Is there a simpler option I'm missing? Something like mTLS with per-user certs, or just accepting static tokens with a short expiry and good auditing? For anyone who went the secret-manager route: did rotation actually work invisibly, or did you get outages from clients that cached the value? submitted by /u/SmartWeb2711 to r/devops [link] [comments]
by SortingYourHosting
For transparency: I run a small UK hosting company, where WHMCS is our current billing/provisioning system. Last week I moved WHMCS 8.13.1 from a CloudLinux OS 9 box running Plesk, to a fresh Debian 13 box. Mostly as we had WHMCS on a shared host during our early days and now it needed its own space. Two things worth sharing if anyone else hits this: Debian 13 ships with PHP 8.4 by default whereas WHMCS wants 8.3 so you need to use the Ondrej Sury repository and not the base repository. ionCube loader has to be placed manually and the ini file renamed to 00-ioncube.ini rather than the default 20-ioncube.ini otherwise if it loads after OPcache, WHMCS breaks very unhelpfully with no clear errors pointing at the fault. Also worth flagging, we had stale absolute paths from the old server that lived in multiple places. We had some in configuration.php, the storage settings gui, and the tblconfiguration SQL table and also in config.php that took a while to track down. Also don't forget to redo the cron! Posting in the hopes it saves someone else an afternoon! submitted by /u/SortingYourHosting to r/sysadmin [link] [comments]