CommPulse

CommPulse

1160 parked Settings

The cross-site community pulse: gold-layer posts + comment threads read live from the Communication Hub, ranked by importance. Turn a post into Discord / LinkedIn / X.

redditFinOpsimportance 0.45View on Reddit

Title vs the work I actually do

by Responsible_Put_2876

Hi and Thank you for reading this. I came into finance sideways. I started in technical support at a SaaS company and gradually became the person handling billing technical escalations: Stripe API integration issues (subscriptions, stripe Connect). That turned into owning bigger pieces, like enabling (using stripe features) revenue teams with some initiatives that didn’t have support on the backend yet. That work got me promoted into what we call a “FinOps” role, and the team wants me as the person driving new initiatives in finance operations. Right now the work is mostly reactive: resolving billing issues, scripting repetitive tasks, and using AI agents where it actually makes sense. There are some projects around ai but nothing related to cloud architecture. Reading this community, I’ve realized my role doesn’t map to what most professionals mean by FinOps. I’m not trying to claim the FinOps title, since I have no DevOps background. But I’d like to hear opinions on how to develop in my current role and what directions are logical. I’d really appreciate any thoughts on this. Thank you. submitted by /u/Responsible_Put_2876 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.45View on Reddit

I got a call from one of our rural customers at this MSP and their coaxial internet was down again. They lose like $1000+ an hour if they're down so they told me they called Verizon, their business cell provider, and asked if they had something. I assume a salesman mentioned it in the past. After one phone call where I said "our networking guy is out today but I'm pretty sure you can plug just about anything into the WAN port on the Fortigate and it'll work" they called back an hour later saying they got an a XC46BE. I've never even heard of that family of devices so I thought this is gonna be a shit day. Luckily for my non-networking specialist self, was easier to set up than a 54g linksys in 2002. Basically it just jumps on and spits out wifi and ethernet. I ran a test on my non-verizon smartphone and got 16mbps to the tower. Then I hooked my laptop up to the device and got 220x40 so fuck net neutrality I guess. It even has a battery and their switches had UPSes so that's a bit interesting. I slapped it into the Fortigate and tada, everyone's back online...with about 10mbps and 500ms ping time. I assumed all their Outlooks were syncing at once or something but they told her the device can do about 20 people. They have about 10 highly active computers. Pretty unimpressive for $350! But they intended to return it when the outage was fixed and their rep said that was fine. The very millisecond I walked out the door, the ISP truck showed up and started messing with the box on the front lawn. Awesome use of my time. But I keep hearing about these magical devices that can switch over to 5G and we do have WAN1 and WAN2 on the Fortigate so they're considering keeping it and programming in a switchover of some sort. I assume they do that. Anyone have one that doesn't suck? Because this one impressed me until it was actually in use. I saw one at a trade show years ago that was a UPS + 5G modem. That sounded kinda neat but so did this Verizon device until it performed poorly. submitted by /u/CeC-P to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.45View on Reddit

I've had multiple computers in our network failing every single Intune install immediately. Like, it fails to even download. The devices are obviously connected to the network, as they are getting an Intune sync (when sent from the Intune portal or requested via Company Portal) but then continuous notifications about apps failing. My gut is telling me that there is some kind of malware that is somehow allowing the Intune sync to go through, but blocking the download traffic. I know it's not a network issue, because it's only certain computers doing this, and even at the same physical site some are fine while others are broken. And the computers can get out to the web, it seems to be just certain traffic, like this, that is not going through. The IntuneManagementExtension.log has this: [Win32App][Win32AppDownloadExecutor] Execution completed with action status: Failed, enforcement state: InProgressPendingManagedInstaller, error code: , and download running in background: False. ... [SendWebRequestInternal] Sending network request... Current proxy is https://agents.msua06.manage.microsoft.com/TrafficGateway/TrafficRoutingService/SideCar/StatelessSideCarGatewayService/SideCarGatewaySessions('955b27ad-37ba-4622-9b92-00a8f01e1737')%3Fapi-version=1.6 [SendWebRequestInternal] Succeeded, client-request-id: 5d90a1b0-f063-40b3-ab88-d9b2d17847eb, AfdRef: Found 1 MDM certificates from Local Computer Store. [Win32App][Win32AppDownloadExecutor] Execution completed with action status: Failed, enforcement state: InProgressPendingManagedInstaller, error code: , and download running in background: False. Checking throttle setting Successfully updated throttling info. workload AgentCheckIn, currentCnt = 5 Finish throttle checking. It's weird that the error code is just empty. And that line is just continuously there, appearing about every 2-3 seconds as it continues to try to pull more apps. Any ideas on how I can troubleshoot this? submitted by /u/havens1515 to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.45View on Reddit

Today we had 3 AVD multisession hosts that went through the Live Migration process at different times of the day (just learned about this today so go easy on me) - https://learn.microsoft.com/en-us/azure/virtual-machines/maintenance-and-updates#live-migration I don't see any errors of the live migration failing, as its showing succeed, but my users are getting kicked off then unable to connect. The error I see when a user tries to connect is: ConnectionFailedUserHasValidSessionButRdshIsUnhealthy So I checked the host and its in a stuck/updating state. PS C:\Users\bob> (Get-AzVM -ResourceGroupName $rg -Name $hostname -Status).VMAgent.Statuses Code : ProvisioningState/Unavailable Level : Warning DisplayStatus : Not Ready Message : VM Agent is unresponsive. Time : 8/4/2026 9:51:20 PM And PS C:\Users\bob> (Get-AzVM -ResourceGroupName $rg -Name $hostname -Status).Statuses | ft Code,DisplayStatus -Auto Code | DisplayStatus ProvisioningState/updating | Updating PowerState/running | VM running Host information: West-US 2 E8s_v6 Windows 11 multisession NERDIO managed I tried connecting through NERDIO Console Connect, RDP IP and Hostname, and then tried running Azure Run Commands, all failed. Was trying to get in to try and fix the agent by restarting it or reinstalling in hopes to fix it. I put the host in Drain Mode, but users were still getting directed to that host if they had a session already there (according to google that's intended). So my only fix was to restart the hosts. Looking for ideas on how to prevent this and or a simple fix when it happens. submitted by /u/SendMe_YourPasswords to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.45View on Reddit

Single Proxmox host today, budget approved to grow to two nodes. A proper Dell/HPE storage array is far out of reach at this budget — but a business-class NAS is not, and that's exactly my question: **would a business NAS actually be enough for us, or am I fooling myself?** Small manufacturing company, ~50 workstations, I'm the only IT person. ## Current host Lenovo ThinkSystem SR650 - 1× Xeon Silver 4210R (10c/20t). Second socket empty. - 32 GB RAM in a single DIMM. 23 slots free. - VM datastore: RAID1 of 2× 2.4TB 10K SAS HDD → 2.2 TB, **94% full**. - Boot/local: RAID1 of 2× 960GB SATA SSD → 893 GiB, ~30% used. - ThinkSystem RAID 730-8i (hardware RAID, no proper JBOD passthrough). - Intel X722 LOM, 4 ports. No 10GbE add-in card, but free PCIe slot. - Core switch is 48-port gigabit with 10G SFP+ uplinks available. - Proxmox VE 8.4.14. ## Workload | Role | vCPU | RAM | |---|---|---| | AD DS + DNS | 4 | 12 GB | | Zabbix + Grafana (LXC) | 4 | 4 GB | | Internal web app | 4 | 2 GB | | Quoting web app | 2 | 4 GB | | UniFi controller | 2 | 2 GB | | GLPI (Docker) | 2 | 2 GB | | API gateway | 1 | 2 GB | **Allocated: 19 of 20 vCPU, 28 of 32 GB RAM.** **Actually used: under 5% CPU, ~17 GB RAM.** A second DC and the SIEM live outside this host. **Veeam runs on its own separate machine and backups also land offsite**, so backup does not depend on this host or on whatever storage we buy. **Coming soon:** internal ERP — web app plus a **MySQL** database. --- ## Option 1 — Two full nodes, local storage, ZFS replication - **New node:** better CPU than current, max 16 cores, 64 GB RAM, 2× 960GB SSD + 3× 3.2TB SAS for capacity. - **Current node:** RAM upgrade, HBA to replace hardware RAID. - ZFS both sides, 10GbE direct link, external QDevice, async replication (~15 min). Two independent copies of the data. Local disk latency. But async replication means up to 15 minutes lost on failover, and I have to swap the RAID controller for an HBA on the existing box. ## Option 2 — Business NAS as shared storage + thin node - **New node:** 2× SSD purely for Proxmox itself, no VM storage. Budget goes into RAM instead of disks. - **Current node:** RAM 32 → 64 GB. Keep the existing controller — **no HBA swap, no ZFS to design.** - **Business-class NAS** holds all VM disks, 10GbE to both nodes. Both nodes mount shared storage → live migration and HA with no RPO gap. Clears the 94% problem on day one. Simpler to build. But one NAS, one controller, and if it dies the whole virtual environment is down until a replacement arrives. ## Budget Up to **R$100,000 ≈ US$19,400** total, covering the new node and either the disks or the NAS. This is Brazil — import duties and local margins mean enterprise hardware here lands well above US list, so that number buys maybe half of what it would in the US. --- ## Questions **The main one: is a business-class NAS genuinely adequate as primary storage here?** Our VMs currently run off two 10K spinning disks in RAID1, so almost anything is an improvement on paper. But "business NAS" covers everything from a 4-bay desktop box to a rackmount unit with redundant PSUs and all-flash. What tier do I need to insist on, and what specs are non-negotiable? **MySQL over the network.** Is a small ERP database the workload where NAS-backed storage falls apart, or is that overblown at this scale? Would you keep the DB on the node's local SSD even in Option 2? **iSCSI + LVM-thin or NFS + qcow2?** iSCSI performs better but I lose snapshots; NFS keeps snapshots but adds latency. Which are you running in production and what do you regret? **Option 2 lets me skip the HBA swap and skip designing ZFS entirely.** Is that a real advantage, or am I trading a solved problem for a worse one? **In Option 1, is 3× SAS HDD sane for a VM datastore in 2026,** or should the whole VM pool be SSD and the spinning disks reserved for documents and archives? **Given the budget and this workload, which would you build?** We already have Veeam on a separate machine with offsite copies, so the question is less about losing data and more about how long we'd be down. Happy to answer questions about the workload. submitted by /u/gianheller to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditawsimportance 0.45View on Reddit

Did a cleanup pass this month after the bill crept up again and it was grim. Over 1k a month going to stuff nothing was using. The usual suspects are unattached EBS volumes from instances we killed months ago, a pile of snapshots from nonexistent volumes, NAT gateways 3 of them just idle in a dev account at 32 bucks a month each and a couple of load balancers with no targets. There was also an elastic IP quietly billing by the hr since AWS started charging for those. Worse than last year, a chunk of it traced back to our own agents. Devs run coding agents that spin up test infra to try something and the agent never tears it down, teardown isn't in the happy path. So every abandoned experiment leaves a little orphaned tail nobody's watching because it's 20 bucks here and forty there til a year of it adds up. Tagging would catch some of this which of course it isn't and the untagged stuff is the orphaned stuff because it got made in a hurry. Cost Explorer shows me the number, never the owner. This is the boring waste that never trips an alarm, it just quietly rents space in your bill forever. submitted by /u/Chris-Hart_232 to r/aws [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.45View on Reddit

I’m trying to understand the best way to investigate a sudden increase in Azure spending. Beyond Azure Cost Management alerts, what tools or practices do you use to identify which resources or workloads are responsible for unexpected cost changes? I’m particularly interested in approaches that work well in environments with multiple subscriptions and resource groups. submitted by /u/DianeAtkinson to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.45View on Reddit

Everyone has been waiting for the Azure price increase tied to the memory shortage. My two cents: it's already here. Normally a generational compute update — v1 through to v5 — costs the same per hour and gives you a performance bump - free of charge.. It also helps Microsoft exit their old End Of Life hardware. Everybody wins. You move v2 to v3 on a 16 vCPU box and the rate is more or less identical, right the way up to v5. That's over. From v6 onwards there's a fundamental shift. The generational upgrade is no longer free. v1 to v5: same price. v6: roughly +10%. v7: roughly +35%. Read that again, because it changes how you have to think about your estate. EOL now exists in the cloud. Not as a migration exercise — as a cost event. Previously, hardware retirement was Microsoft's problem. They wanted you off the old fleet, so they made the move painless and you got free performance out of it. Now the retirement notice comes with a bill attached, and you have no route to decline it. Sub-v5 capacity is already constrained. Once the capacity pressure tightens further, "stay where you are" stops being an option. So combine three things: → Generational moves are now priced increases, not neutral swaps → Capacity constraints on older SKUs push you up the generations whether you budgeted for it or not → Nobody's three-year plan has a compounding uplift modelled into it That's a ticking time bomb. Your Azure VMs are now going to get generationally more expensive by default. Not because anyone announced a price rise — because the escalator only goes one way and you're standing on it. submitted by /u/AcceptablePicture329 to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditgooglecloudimportance 0.45View on Reddit

Wondering if anyone else is experiencing a slowdown with Firestore query in the last few weeks, which has really started to tick up in the last few days for us. We’re seeing recurring waves of timeouts across multiple unrelated Google Cloud projects and I’m trying to figure out whether anyone else is seeing something similar. Setup: - Firebase Functions Gen 2 / Cloud Run - Node.js Admin SDK - Firestore Native mode - Region: us-central1 - Multiple separate projects affected - Request timeout varies by service, usually 30s or 60s (depending on the function timeout we've set) During a wave, unrelated HTTP routes start returning 504s that approach the function's configured timeout (e.g. 29.997s for a 30s timeout). The failed requests are not tied to one endpoint or one query shape. Logs show Firestore operations continuing after the Cloud Run request has already timed out. Examples from one incident: - `devices.read` direct document read: 116,019ms - `leads.read` direct document read: 57,487ms - `postalCodes.read` direct document read: 33,093ms - `cache.list limit=1`: 61,590ms - `preferences.read` direct document read: 36,324ms - `leads.list limit=500`: 79,723ms The `postalCodes.read` example is especially confusing because that document is tiny. But honestly all of these documents are a few kb at the most. So this doesn’t look like just a large document, missing index, or bad query issue. Other observations: - Failures often cluster on one Cloud Run instance/revision instance ID. - Other instances may continue serving traffic normally. - The request hits the Cloud Run timeout, but the Firestore operation later logs completion. - We’ve seen this on three separate projects. - We’ve reduced external API calls and removed some broad Firestore scans, but the issue still recurs. - We briefly tried `preferRest: true`; it may have helped temporarily but did not clearly eliminate the issue. Has anyone else seen Firestore Admin SDK calls intermittently hang like this from Cloud Run / Firebase Functions Gen 2? Is there a known issue with Firestore gRPC/transport connections getting unhealthy per instance? Are there recommended client-side mitigations besides lowering concurrency, recycling instances, or reducing reads (i.e. increasing cache usage)? Is there a good way to prove this is client transport/backend behavior versus application query pressure? I'm trying to understand whether this is a known pattern and what people have done to debug or mitigate it. submitted by /u/ehed to r/googlecloud [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
devtofeed/tag/devopsimportance 0.45View on devto

The Paradigm Shift in Enterprise Attack Surfaces The integration of autonomous artificial intelligence agents and dynamic execution layers into enterprise software has fundamentally altered the modern corporate attack surface. When an organization deploys an AI platform capable of generating and executing code on the fly, it effectively introduces an internal multi-tenant execution environment within its perimeter. The security of this architecture relies entirely on the absolute integrity of the boundary separating untrusted, AI-generated code from the underlying host operating system. When this boundary fails, the consequences are immediate: remote code execution (RCE) with the privileges of the orchestration service. I observe that security teams frequently treat AI platforms as standard web applications, focusing their defense-in-depth strategies on traditional web vulnerabilities like SQL injection or Cross-Site Scripting (XSS). However, the deployment of LLM-driven code interpreters introduces a fundamentally different threat model. In this paradigm, the application is designed to execute arbitrary code by design. The primary security control is no longer input sanitization alone, but robust, low-level virtualization and process isolation. In this analysis, I examine the structural mechanics of sandbox escape vulnerabilities within enterprise AI platforms, using the architectural patterns of CVE-2026-6875 as a reference model. I will dissect how these boundaries are bypassed, evaluate the operational trade-offs of various isolation technologies, and provide a concrete blueprint for hardening your runtime infrastructure. My objective is to move beyond superficial patching advice and equip engineering leaders with the technical depth required to design resilient execution environments. An in-depth technical analysis of CVE-2026-6875, a critical CVSS 9.5 sandbox escape vulnerability in the ServiceNow AI Platform under active exploitation. Learn the mechanics of AI execution layer esc 🤖 Architectural Analysis of AI Execution Layers To understand how a sandbox escape occurs, one must first analyze the typical architecture of an enterprise AI execution engine. These systems generally consist of three primary components: the orchestration layer, the communication bridge, and the isolated guest runtime. The Orchestration Layer This component runs on the host system, often with high privileges. It interfaces with the broader enterprise application, manages user sessions, and coordinates with the LLM. When the LLM determines that a task requires code execution (such as data analysis, mathematical computation, or file parsing), the orchestration layer receives the generated code block and prepares the execution environment. The Communication Bridge Because the host must pass code into the sandbox and retrieve the execution output, a communication channel must exist. This is typically implemented via Unix domain sockets, local loopback TCP connections, or shared memory segments. A daemon running inside the sandbox listens on this channel, receives the code, executes it via a local interpreter (such as Python or Node.js), and returns the stdout, stderr, and generated files. The Isolated Guest Runtime This is the sandbox itself. In many standard deployments, this is a lightweight container managed by runc, Docker, or containerd. The isolation relies on standard Linux kernel features: namespaces (to isolate processes, network interfaces, mount points, and IPC), control groups (cgroups, to limit CPU, memory, and I/O), and seccomp filters (to restrict available system calls). I must emphasize that this architecture contains an inherent structural tension. The AI agent requires access to libraries, packages, and sometimes external APIs to perform useful work. However, every capability granted to the guest runtime increases the available attack surface. If the guest runtime shares the host operating system's kernel—as is the case with standard containerization—any vulnerability in the kernel's system call interface or any misconfiguration in the orchestration layer can be leveraged to escape the container. Deconstructing the Escape Mechanics My analysis of sandbox escapes reveals that failures rarely occur within the isolated guest runtime itself. Instead, the breakdown typically occurs at the interface between the host and the guest, or through the exploitation of shared kernel resources. I have categorized the primary escape vectors into four distinct operational patterns. 1. Kernel Interface Exploitation and Shared Syscalls When standard containers are used for sandboxing, the guest processes execute directly on the host kernel. If an attacker can execute arbitrary code inside the guest, they can interact directly with the host kernel via system calls. If a local privilege escalation (LPE) vulnerability exists in the host kernel (for example, in memory management or network namespaces), the attacker can exploit it from within the container to gain root privileges on the host, subsequently breaking out of the container namespaces. 2. Orchestration Agent Command Injection This is a common vulnerability pattern in custom-built AI platforms. The orchestration layer on the host often uses command-line utilities to manage the lifecycle of the sandbox (e.g., spawning containers via docker run or executing commands inside them via docker exec ). If the orchestration layer fails to properly sanitize parameters passed to these commands—such as environment variables, volume mount paths, or container names—an attacker can inject shell metacharacters. This results in command execution on the host system, completely bypassing the sandbox. 3. Socket and IPC Hijacking To monitor container health or manage files, developers sometimes mount the host's container runtime socket (such as /var/run/docker.sock ) inside the sandbox. This is an architectural anti-pattern of the highest severity. If an attacker gains code execution inside a sandbox with access to this socket, they can issue API commands to the host's container daemon to spawn a new container. This new container can be configured with host namespaces, host network access, and the host's root directory mounted as a volume, yielding immediate and total control over the host system. 4. Path Traversal and Shared Volume Manipulation To facilitate file input and output, the host must share a directory with the sandbox. This is typically achieved via bind mounts. If the host-side application processes files written to this shared directory without rigorous validation, several vulnerabilities can emerge. For example, if the guest runtime creates a symbolic link pointing to a sensitive host file (such as /etc/shadow or /root/.ssh/authorized_keys ) within the shared directory, and the host application reads or writes to that link without verifying that it resolves within the allowed boundary, the host will inadvertently read or overwrite its own system files. Operational Trade-offs of Isolation Technologies When designing a remediation strategy, engineering leaders must choose an isolation technology that balances security, performance, and operational complexity. I have evaluated the three primary paradigms currently used in production environments. Standard Containers (runc / Docker / Kubernetes) Mechanism: OS-level virtualization sharing the host kernel, isolated via namespaces and cgroups. Security Posture: Low. The shared kernel design means any kernel vulnerability can lead to a complete host compromise. It is highly susceptible to configuration drift and misconfigurations. Performance: Excellent. Near-zero virtualization overhead; sub-second startup times. Operational Complexity: Low. Standard tooling and deep ecosystem integration. User-Space Kernels (gVisor) Mechanism: A runc-compatible container runtime that intercepts system calls from the guest application and filters them through a user-space kernel (written in Go) before passing a limited subset to the host kernel. Security Posture: Medium-High. It drastically reduces the host kernel attack surface by blocking direct system calls. Even if a guest process attempts to exploit a kernel vulnerability, the system call is intercepted and handled safely in user space. Performance: Moderate. The system call interception introduces latency, which can impact I/O-heavy or system-call-intensive workloads. Operational Complexity: Moderate. It integrates with existing container orchestrators like Kubernetes but requires specific runtime class configurations. My Recommendation: I recommend gVisor as an excellent compromise for organizations with existing Kubernetes infrastructure that cannot easily transition to hardware virtualization. MicroVMs (AWS Firecracker / Kata Containers) Mechanism: Hardware-assisted virtualization using the Linux Kernel-based Virtual Machine (KVM) hypervisor to launch extremely lightweight, minimalist virtual machines with their own dedicated kernels. Security Posture: High. The boundary is enforced at the hardware level. A sandbox escape requires exploiting the hypervisor itself, which is a significantly smaller and more secure interface than the Linux kernel system call interface. Performance: High. Startup times are measured in milliseconds (typically under 100ms), and memory overhead is minimal compared to traditional virtual machines. Operational Complexity: High. Requires bare-metal instances or nested virtualization support in cloud environments. It also requires specialized orchestration tooling. My Recommendation: For dedicated AI execution engines handling highly untrusted code, I consider microVMs to be the gold standard. The security benefits far outweigh the initial setup complexity. ⚙️ Implementation: A Secure Execution Wrapper To illustrate the practical application of these hardening principles, I have designed a robust Python execution wrapper. This implementation demonstrates how to enforce strict resource constraints, drop privileges, and isolate execution using standard Linux system controls before running untrusted code. I must emphasize that while this wrapper significantly hardens standard process execution, it should be deployed inside an isolated container or microVM to achieve true defense-in-depth. import os import sys import pwd import grp import resource import subprocess import tempfile import shutil from pathlib import Path def enforce_sandbox_limits(uid: int, gid: int, max_cpu_seconds: int = 5, max_memory_bytes: int = 128 * 1024 * 1024): """ Configures process limits and drops privileges to an unprivileged user. This function must run in the child process before executing untrusted code. """ # 1. Establish strict resource limits (cgroups equivalent at process level) # Limit CPU time to prevent infinite loops and denial of service resource.setrlimit(resource.RLIMIT_CPU, (max_cpu_seconds, max_cpu_seconds)) # Limit virtual memory allocation to prevent memory exhaustion resource.setrlimit(resource.RLIMIT_AS, (max_memory_bytes, max_memory_bytes)) # Limit file creation size to prevent disk filling resource.setrlimit(resource.RLIMIT_FSIZE, (1024 * 1024, 1024 * 1024)) # 1 MB # Limit number of processes to prevent fork bombs resource.setrlimit(resource.RLIMIT_NPROC, (20, 20)) # Disable core dumps to prevent sensitive data leakage resource.setrlimit(resource.RLIMIT_CORE, (0, 0)) # 2. Drop privileges to the designated unprivileged user try: os.setgroups([]) # Clear supplementary groups os.setgid(gid) os.setuid(uid) # Ensure we cannot regain root privileges os.environ["USER"] = pwd.getpwuid(uid).pw_name os.environ["HOME"] = pwd.getpwuid(uid).pw_dir except Exception as e: sys.stderr.write(f"[FATAL] Failed to drop privileges: {str(e)}\n") sys.exit(1) def execute_untrusted_code(code_payload: str, sandbox_user: str = "sandbox_worker") -> dict: """ Executes untrusted Python code within a restricted, temporary directory under an unprivileged system account with strict resource constraints. """ # Resolve the unprivileged user credentials try: user_info = pwd.getpwnam(sandbox_user) uid = user_info.pw_uid gid = user_info.pw_gid except KeyError: return { "success": False, "error": f"System user '{sandbox_user}' does not exist. Aborting for safety." } # Create an isolated temporary directory for execution temp_dir = tempfile.mkdtemp(prefix="ai_sandbox_") temp_path = Path(temp_dir) script_path = temp_path / "payload.py" try: # Write the payload to the temporary directory script_path.write_text(code_payload, encoding="utf-8") # Adjust ownership of the directory and file to the unprivileged user os.chown(temp_dir, uid, gid) os.chown(str(script_path), uid, gid) # Restrict permissions: only the owner can read/write/execute os.chmod(temp_dir, 0o700) os.chmod(str(script_path), 0o500) # Read and execute only for the worker # Execute the untrusted script in a isolated subprocess process = subprocess.run( [sys.executable, str(script_path)], preexec_fn=lambda: enforce_sandbox_limits(uid, gid), capture_output=True, text=True, timeout=10, # Hard wall-clock timeout cwd=temp_dir ) return { "success": True, "return_code": process.returncode, "stdout": process.stdout, "stderr": process.stderr } except subprocess.TimeoutExpired as e: return { "success": False, "error": f"Execution exceeded maximum wall-clock time limit of {e.timeout} seconds.", "stdout": e.stdout.decode() if e.stdout else "", "stderr": e.stderr.decode() if e.stderr else "" } except Exception as e: return { "success": False, "error": f"Internal execution failure: {str(e)}" } finally: # Securely clean up the execution directory try: shutil.rmtree(temp_dir) except Exception as cleanup_error: sys.stderr.write(f"[ERROR] Failed to clean up sandbox directory {temp_dir}: {str(cleanup_error)}\n") Immediate Incident Response and Remediation Playbook If you are running enterprise AI platforms that execute dynamic code, you must assume that your systems are targeted. I recommend executing the following response protocol immediately to identify potential compromises and secure your infrastructure. Step 1: Asset Discovery and Mapping You must identify every instance of the AI execution engine within your environment. This includes production servers, development environments, staging environments, and any local testing instances. Attackers frequently target unmonitored staging environments to establish an initial foothold, then pivot laterally into production networks. 🔐 Step 2: Log Analysis and Threat Hunting Do not rely solely on automated alerts. I recommend performing a manual, structured audit of your system and application logs, looking back at least 90 days. Focus your investigation on the following indicators: Process Creation Logs: Audit your host operating system logs (such as auditd or Sysmon) for anomalous processes spawned by the AI service user. Look specifically for shells ( /bin/sh , /bin/bash , /bin/zsh ), network utilities ( curl , wget , nc , socat ), or compiler tools ( gcc , make ). Network Flow Logs: Analyze outbound connection records from your AI worker nodes. Any connection initiated by an AI worker to an external IP address—especially those not explicitly whitelisted—must be treated as highly suspicious. Pay close attention to connections targeting cloud metadata services (e.g., 169.254.169.254 ). File Integrity Monitoring: Check for modifications to critical system files, SSH configuration directories ( ~/.ssh/ ), cron jobs, or systemd service files on the host operating system running the AI platform. Step 3: Network Segmentation and Virtual Patching If an official patch cannot be applied immediately due to operational constraints, you must implement strict network segmentation. Isolate the AI execution hosts. Block all outbound internet access from these hosts, and restrict inbound traffic to authenticated, internal corporate networks. If the AI engine requires external data, route those requests through a secure, validating proxy server on the host. Long-Term Hardening Framework To move beyond reactive patching and build a truly resilient AI infrastructure, engineering teams must adopt a zero-trust model for code execution. I have compiled a structured checklist of critical controls that you should audit and implement immediately. Control Domain Security Requirement Implementation Verification Process Isolation AI execution runtimes must run under dedicated, non-root, unprivileged system accounts. Verify that the UID of the running process inside the container is not 0. Resource Constraints Hard limits must be enforced on CPU, memory, process count, and disk write sizes. Verify cgroup configurations and process limits ( ulimit ) on the host. Network Isolation Outbound network access from the sandbox to internal networks and cloud metadata APIs must be blocked. Attempt to curl 169.254.169.254 or an internal IP from within the sandbox; it must fail. Filesystem Security The root filesystem of the sandbox must be mounted as read-only. Attempt to write to /usr , /bin , or /etc from within the sandbox; it must fail. Boundary Validation All data passed between the host and the sandbox must be strictly validated against a schema. Ensure that the communication bridge does not accept raw shell commands or unvalidated file paths. Ephemeral Lifecycles Sandbox environments must be destroyed and recreated after every execution task. Verify that no state or files persist between separate execution requests. 🎯 Conclusion The emergence of sandbox escape vulnerabilities in enterprise AI platforms is a predictable consequence of the rapid integration of dynamic execution layers. When we design systems that allow models to write and execute code, we must abandon the assumption that the code is benign. We must design our architectures with the fundamental assumption that the sandbox will be compromised. By transitioning from shared-kernel container isolation to hardware-level microVM virtualization, enforcing strict system call filtering, dropping process privileges, and implementing zero-trust network policies, you can ensure that a sandbox escape remains an isolated event rather than an enterprise-wide catastrophe. Security must not be treated as a feature to be added later; it must be the architectural foundation upon which your AI infrastructure is built. 🔗 Originally published on ixuvo.com

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.45View on Reddit

​ I recently opened an AWS Support case regarding an unexpected billing issue. The case was created successfully, but its status is currently showing as “Unassigned.” Case number: 178740096600724 It has been 24 hr since I created the case. Has anyone experienced this before? How long did it take for your AWS Billing Support case to get assigned to an agent? I’m mainly looking to understand whether I should just wait or take any further action. Thanks! submitted by /u/Separate_Break8620 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.45View on Reddit

I have 2 Standalone HPE Server where just Windows Server 2022 is running and unfortunately a bios update corrupted bios and made HPE Server unable to boot. This has unfortunately been caused by a power outage during planned maintenance bios update. So I ordered a CH341a programmer and flashed the stock bios from hpe to this mainboard. Both systems booted fine again however users couldn't connect to file shares of one of them anymore due to duplicated uuids and only one of the system can be in the ad domain at the same time. Also mac addresses are the same on both systems so I had to set a static one for one of them in device managers nic driver. Is there any way to fixx the broken bios update? submitted by /u/luky90 to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.45View on Reddit

so this happened our old phone system was a mess. on-prem PBX from like 2015. constantly dropping calls, hard to scale, and every time something broke I had to drive to the office to fix it. I hate driving to the office. I pitched moving everything to Azure. compute, storage, the whole thing. took me like 3 weeks to build the business case. showed him cost projections, uptime improvements, scalability. he finally said yes. we went with phone system for the actual calling layer and built the rest on Azure. recordings go to blob storage, analytics run on functions, the whole stack. honestly it's been solid. we had one hiccup with network security groups blocking SIP traffic but that was my fault for not configuring it properly. now whenever someone asks about the phone system I'm like it's on Azure and they nod like I'm some kind of genius. I'm not. I just read a lot of documentation. the best part is I haven't had to drive to the office in 4 months. that alone is worth the migration. anyone else running voice workloads on Azure. any tips for cost optimization because I'm terrified of the next bill submitted by /u/Ramsesthrowaway to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.45View on Reddit

Disclosure: I maintain this (free, open source, self-hosted). If your Kubernetes pods request a lot more CPU/memory than they use, you are paying for idle capacity. Attune watches real usage and right-sizes those requests, often without restarting pods (in-place resize on modern Kubernetes). Repo: https://github.com/attune-io/attune Docs: https://attune-io.github.io/attune/ Requirement: usage metrics in the cluster (Prometheus is the usual case; Datadog/CloudWatch also work). Without metrics there is nothing to right-size from. If underuse is real for you, what is the barrier to starting and saving money? Do not trust automation on prod Already use something else No metrics / install friction Hard to prove savings in $ Change management / security What would block you most? submitted by /u/Mobidic69 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditdevopsimportance 0.44View on Reddit

Maybe we need to change the way we handle access process for our developers or users. There was a production outage but luckily it wasn't revenue impacting. I had to help a developer by accessing their application on an ec2 instance. Our team have access to any servers in production. Our developers only have access to our dev and stage environments. I am not familiar with their application. So basically, I was just executing commands that he was giving me. It was the most degrading role I have experienced, HAHAHA! I'm thinking that when there are production outages, the application owners should be given temporary access so they can debug their applications. It will be quicker. It took us almost 5 hours! I was just copying and pasting commands and outputs. On the unix history command recalls everything. I don't recall any, HAHAHA! So what is your process? submitted by /u/Oxffff0000 to r/devops [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.44View on Reddit

Most teams I talked to don't have a CI/CD pipeline checking cost or governance policy before terraform apply, they just run it locally. That's what guard is for: cloudcosttree guard -- terraform apply. To be clear, this isn't a simulation, it's your real terraform apply. guard never runs one on its own, it only wraps the exact command you were already going to run, checks the plan against your policies, then applies that same saved plan, so there's no gap between what got checked and what got deployed. Default behavior is warn-only, it prints violations but still applies. --block opts into actually stopping the apply on a real violation. A false positive blocking a real deploy is worse than one showing up in a report, so blocking is never the default. submitted by /u/Independent-Ease-609 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.44View on Reddit

I'm still seeing the old GPT pricing in Foundry Monitoring and in my Azure cost management for Standard Global as of today, does anyone know when they will fix this? It's been 12 days now... Support isn't helpful and doesn't have any clue it seems. submitted by /u/googleaddreddit to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditawsimportance 0.44View on Reddit

Context: I'm running a public API on API Gateway + Lambda, with DynamoDB behind it. Endpoints have very different backend costs, cheap reads vs. a couple of routes that kick off heavier aggregation work. Currently using API Gateway usage plans with a single throttle limit per API key, applied flat across all routes. The problem: usage plans throttle by requests/second regardless of which route is hit, so a client hammering cheap GETs eats the same budget as one calling the expensive routes, and there's no way (as far as I can find) to weight individual routes differently within a single usage plan without splitting them into separate API Gateway stages/plans per cost tier, which gets awkward to manage as the number of "cost classes" grows. What I've looked at so far: Per-stage/per-plan splitting: works, but means maintaining N usage plans and N sets of API keys per client if a client needs access to routes at more than one cost tier. Custom Lambda authorizer + DynamoDB counter: doing weighted token-bucket logic myself (consume different token amounts per route, check/decrement atomically via DynamoDB conditional writes), seems doable but adds a DynamoDB read/write on every request just for the rate-limit check, plus I'd be reimplementing throttling that API Gateway mostly already does for free. Briefly looked at whether Lambda reserved/provisioned concurrency per function could act as an implicit cost-based limiter (route the expensive endpoint through its own function with tighter concurrency), but that limits total throughput, not per-client fairness. Has anyone actually shipped weighted/cost-based rate limiting on top of API Gateway usage plans, or does everyone end up rolling their own with a Lambda authorizer + DynamoDB/ElastiCache counter once costs diverge enough between routes? And if you rolled your own, did you keep API Gateway's built-in throttling as a coarse backstop on top of it, or drop it entirely in favor of the custom logic? submitted by /u/ClickOk5811 to r/aws [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.44View on Reddit

I need some help figuring out if this is a highly targeted scam or a legitimate invoice. I recently received an email containing a PDF invoice for an "Unreturned Advance Exchange Fee". The sender email is ⁠ [email protected] ⁠. Here is what is really tripping me up and making me second-guess: The details of the item on the invoice (a Surface Laptop) and the serial number of the returned item is correct. My personal and organization details listed in the "Bill To" and "Ship To" sections are 100% correct. The invoice claims to be from "MICROSOFT PTY LIMITED" based in North Sydney, Australia, and the total amount is listed in AUD. **The Red Flag:** The "Remit to Bank" section instructs me to send payment to "BANK OF AMERICA". I am based in the Asia Pacific region, so seeing Bank of America as the payment destination feels incredibly suspicious, even though Microsoft is a US-based company. Has anyone else dealt with this before? Is it normal for Microsoft's APAC/Australian branches to use Bank of America for direct wire transfers, or is this just a very sophisticated, highly personalized spoofing attempt using an email address like **⁠ [email protected] **⁠? Customer Success Manager from MS hasn’t replied in a week where I asked if this is legit and I should reply to it. Also btw device was returned within directed time frame. Any advice would be greatly appreciated! submitted by /u/laleric to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.44View on Reddit

terraform plan validates syntax and state, it has no opinion on whether the instance size you picked is actually right for the workload. Ran a single-resource RDS example through CloudCostTree: db.t3.large, declared as-is, comes out to $107.28/month, Cost Score F. Simulating one size down, db.t3.medium, same file, no architecture change: $57.64/month. A real $49.64/mo drop, 46.3%. Nothing about that shows up in a terraform plan diff, you only see it once you actually price the resource. submitted by /u/Independent-Ease-609 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X