CommPulse

CommPulse

1160 parked Settings

The cross-site community pulse: gold-layer posts + comment threads read live from the Communication Hub, ranked by importance. Turn a post into Discord / LinkedIn / X.

redditsysadminimportance 0.45View on Reddit

Single Proxmox host today, budget approved to grow to two nodes. A proper Dell/HPE storage array is far out of reach at this budget — but a business-class NAS is not, and that's exactly my question: **would a business NAS actually be enough for us, or am I fooling myself?** Small manufacturing company, ~50 workstations, I'm the only IT person. ## Current host Lenovo ThinkSystem SR650 - 1× Xeon Silver 4210R (10c/20t). Second socket empty. - 32 GB RAM in a single DIMM. 23 slots free. - VM datastore: RAID1 of 2× 2.4TB 10K SAS HDD → 2.2 TB, **94% full**. - Boot/local: RAID1 of 2× 960GB SATA SSD → 893 GiB, ~30% used. - ThinkSystem RAID 730-8i (hardware RAID, no proper JBOD passthrough). - Intel X722 LOM, 4 ports. No 10GbE add-in card, but free PCIe slot. - Core switch is 48-port gigabit with 10G SFP+ uplinks available. - Proxmox VE 8.4.14. ## Workload | Role | vCPU | RAM | |---|---|---| | AD DS + DNS | 4 | 12 GB | | Zabbix + Grafana (LXC) | 4 | 4 GB | | Internal web app | 4 | 2 GB | | Quoting web app | 2 | 4 GB | | UniFi controller | 2 | 2 GB | | GLPI (Docker) | 2 | 2 GB | | API gateway | 1 | 2 GB | **Allocated: 19 of 20 vCPU, 28 of 32 GB RAM.** **Actually used: under 5% CPU, ~17 GB RAM.** A second DC and the SIEM live outside this host. **Veeam runs on its own separate machine and backups also land offsite**, so backup does not depend on this host or on whatever storage we buy. **Coming soon:** internal ERP — web app plus a **MySQL** database. --- ## Option 1 — Two full nodes, local storage, ZFS replication - **New node:** better CPU than current, max 16 cores, 64 GB RAM, 2× 960GB SSD + 3× 3.2TB SAS for capacity. - **Current node:** RAM upgrade, HBA to replace hardware RAID. - ZFS both sides, 10GbE direct link, external QDevice, async replication (~15 min). Two independent copies of the data. Local disk latency. But async replication means up to 15 minutes lost on failover, and I have to swap the RAID controller for an HBA on the existing box. ## Option 2 — Business NAS as shared storage + thin node - **New node:** 2× SSD purely for Proxmox itself, no VM storage. Budget goes into RAM instead of disks. - **Current node:** RAM 32 → 64 GB. Keep the existing controller — **no HBA swap, no ZFS to design.** - **Business-class NAS** holds all VM disks, 10GbE to both nodes. Both nodes mount shared storage → live migration and HA with no RPO gap. Clears the 94% problem on day one. Simpler to build. But one NAS, one controller, and if it dies the whole virtual environment is down until a replacement arrives. ## Budget Up to **R$100,000 ≈ US$19,400** total, covering the new node and either the disks or the NAS. This is Brazil — import duties and local margins mean enterprise hardware here lands well above US list, so that number buys maybe half of what it would in the US. --- ## Questions **The main one: is a business-class NAS genuinely adequate as primary storage here?** Our VMs currently run off two 10K spinning disks in RAID1, so almost anything is an improvement on paper. But "business NAS" covers everything from a 4-bay desktop box to a rackmount unit with redundant PSUs and all-flash. What tier do I need to insist on, and what specs are non-negotiable? **MySQL over the network.** Is a small ERP database the workload where NAS-backed storage falls apart, or is that overblown at this scale? Would you keep the DB on the node's local SSD even in Option 2? **iSCSI + LVM-thin or NFS + qcow2?** iSCSI performs better but I lose snapshots; NFS keeps snapshots but adds latency. Which are you running in production and what do you regret? **Option 2 lets me skip the HBA swap and skip designing ZFS entirely.** Is that a real advantage, or am I trading a solved problem for a worse one? **In Option 1, is 3× SAS HDD sane for a VM datastore in 2026,** or should the whole VM pool be SSD and the spinning disks reserved for documents and archives? **Given the budget and this workload, which would you build?** We already have Veeam on a separate machine with offsite copies, so the question is less about losing data and more about how long we'd be down. Happy to answer questions about the workload. submitted by /u/gianheller to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditawsimportance 0.45View on Reddit

Did a cleanup pass this month after the bill crept up again and it was grim. Over 1k a month going to stuff nothing was using. The usual suspects are unattached EBS volumes from instances we killed months ago, a pile of snapshots from nonexistent volumes, NAT gateways 3 of them just idle in a dev account at 32 bucks a month each and a couple of load balancers with no targets. There was also an elastic IP quietly billing by the hr since AWS started charging for those. Worse than last year, a chunk of it traced back to our own agents. Devs run coding agents that spin up test infra to try something and the agent never tears it down, teardown isn't in the happy path. So every abandoned experiment leaves a little orphaned tail nobody's watching because it's 20 bucks here and forty there til a year of it adds up. Tagging would catch some of this which of course it isn't and the untagged stuff is the orphaned stuff because it got made in a hurry. Cost Explorer shows me the number, never the owner. This is the boring waste that never trips an alarm, it just quietly rents space in your bill forever. submitted by /u/Chris-Hart_232 to r/aws [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.45View on Reddit

I’m trying to understand the best way to investigate a sudden increase in Azure spending. Beyond Azure Cost Management alerts, what tools or practices do you use to identify which resources or workloads are responsible for unexpected cost changes? I’m particularly interested in approaches that work well in environments with multiple subscriptions and resource groups. submitted by /u/DianeAtkinson to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.45View on Reddit

Everyone has been waiting for the Azure price increase tied to the memory shortage. My two cents: it's already here. Normally a generational compute update — v1 through to v5 — costs the same per hour and gives you a performance bump - free of charge.. It also helps Microsoft exit their old End Of Life hardware. Everybody wins. You move v2 to v3 on a 16 vCPU box and the rate is more or less identical, right the way up to v5. That's over. From v6 onwards there's a fundamental shift. The generational upgrade is no longer free. v1 to v5: same price. v6: roughly +10%. v7: roughly +35%. Read that again, because it changes how you have to think about your estate. EOL now exists in the cloud. Not as a migration exercise — as a cost event. Previously, hardware retirement was Microsoft's problem. They wanted you off the old fleet, so they made the move painless and you got free performance out of it. Now the retirement notice comes with a bill attached, and you have no route to decline it. Sub-v5 capacity is already constrained. Once the capacity pressure tightens further, "stay where you are" stops being an option. So combine three things: → Generational moves are now priced increases, not neutral swaps → Capacity constraints on older SKUs push you up the generations whether you budgeted for it or not → Nobody's three-year plan has a compounding uplift modelled into it That's a ticking time bomb. Your Azure VMs are now going to get generationally more expensive by default. Not because anyone announced a price rise — because the escalator only goes one way and you're standing on it. submitted by /u/AcceptablePicture329 to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditgooglecloudimportance 0.45View on Reddit

Wondering if anyone else is experiencing a slowdown with Firestore query in the last few weeks, which has really started to tick up in the last few days for us. We’re seeing recurring waves of timeouts across multiple unrelated Google Cloud projects and I’m trying to figure out whether anyone else is seeing something similar. Setup: - Firebase Functions Gen 2 / Cloud Run - Node.js Admin SDK - Firestore Native mode - Region: us-central1 - Multiple separate projects affected - Request timeout varies by service, usually 30s or 60s (depending on the function timeout we've set) During a wave, unrelated HTTP routes start returning 504s that approach the function's configured timeout (e.g. 29.997s for a 30s timeout). The failed requests are not tied to one endpoint or one query shape. Logs show Firestore operations continuing after the Cloud Run request has already timed out. Examples from one incident: - `devices.read` direct document read: 116,019ms - `leads.read` direct document read: 57,487ms - `postalCodes.read` direct document read: 33,093ms - `cache.list limit=1`: 61,590ms - `preferences.read` direct document read: 36,324ms - `leads.list limit=500`: 79,723ms The `postalCodes.read` example is especially confusing because that document is tiny. But honestly all of these documents are a few kb at the most. So this doesn’t look like just a large document, missing index, or bad query issue. Other observations: - Failures often cluster on one Cloud Run instance/revision instance ID. - Other instances may continue serving traffic normally. - The request hits the Cloud Run timeout, but the Firestore operation later logs completion. - We’ve seen this on three separate projects. - We’ve reduced external API calls and removed some broad Firestore scans, but the issue still recurs. - We briefly tried `preferRest: true`; it may have helped temporarily but did not clearly eliminate the issue. Has anyone else seen Firestore Admin SDK calls intermittently hang like this from Cloud Run / Firebase Functions Gen 2? Is there a known issue with Firestore gRPC/transport connections getting unhealthy per instance? Are there recommended client-side mitigations besides lowering concurrency, recycling instances, or reducing reads (i.e. increasing cache usage)? Is there a good way to prove this is client transport/backend behavior versus application query pressure? I'm trying to understand whether this is a known pattern and what people have done to debug or mitigate it. submitted by /u/ehed to r/googlecloud [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.45View on Reddit

​ I recently opened an AWS Support case regarding an unexpected billing issue. The case was created successfully, but its status is currently showing as “Unassigned.” Case number: 178740096600724 It has been 24 hr since I created the case. Has anyone experienced this before? How long did it take for your AWS Billing Support case to get assigned to an agent? I’m mainly looking to understand whether I should just wait or take any further action. Thanks! submitted by /u/Separate_Break8620 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.45View on Reddit

I have 2 Standalone HPE Server where just Windows Server 2022 is running and unfortunately a bios update corrupted bios and made HPE Server unable to boot. This has unfortunately been caused by a power outage during planned maintenance bios update. So I ordered a CH341a programmer and flashed the stock bios from hpe to this mainboard. Both systems booted fine again however users couldn't connect to file shares of one of them anymore due to duplicated uuids and only one of the system can be in the ad domain at the same time. Also mac addresses are the same on both systems so I had to set a static one for one of them in device managers nic driver. Is there any way to fixx the broken bios update? submitted by /u/luky90 to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.45View on Reddit

so this happened our old phone system was a mess. on-prem PBX from like 2015. constantly dropping calls, hard to scale, and every time something broke I had to drive to the office to fix it. I hate driving to the office. I pitched moving everything to Azure. compute, storage, the whole thing. took me like 3 weeks to build the business case. showed him cost projections, uptime improvements, scalability. he finally said yes. we went with phone system for the actual calling layer and built the rest on Azure. recordings go to blob storage, analytics run on functions, the whole stack. honestly it's been solid. we had one hiccup with network security groups blocking SIP traffic but that was my fault for not configuring it properly. now whenever someone asks about the phone system I'm like it's on Azure and they nod like I'm some kind of genius. I'm not. I just read a lot of documentation. the best part is I haven't had to drive to the office in 4 months. that alone is worth the migration. anyone else running voice workloads on Azure. any tips for cost optimization because I'm terrified of the next bill submitted by /u/Ramsesthrowaway to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.45View on Reddit

Disclosure: I maintain this (free, open source, self-hosted). If your Kubernetes pods request a lot more CPU/memory than they use, you are paying for idle capacity. Attune watches real usage and right-sizes those requests, often without restarting pods (in-place resize on modern Kubernetes). Repo: https://github.com/attune-io/attune Docs: https://attune-io.github.io/attune/ Requirement: usage metrics in the cluster (Prometheus is the usual case; Datadog/CloudWatch also work). Without metrics there is nothing to right-size from. If underuse is real for you, what is the barrier to starting and saving money? Do not trust automation on prod Already use something else No metrics / install friction Hard to prove savings in $ Change management / security What would block you most? submitted by /u/Mobidic69 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditdevopsimportance 0.44View on Reddit

Maybe we need to change the way we handle access process for our developers or users. There was a production outage but luckily it wasn't revenue impacting. I had to help a developer by accessing their application on an ec2 instance. Our team have access to any servers in production. Our developers only have access to our dev and stage environments. I am not familiar with their application. So basically, I was just executing commands that he was giving me. It was the most degrading role I have experienced, HAHAHA! I'm thinking that when there are production outages, the application owners should be given temporary access so they can debug their applications. It will be quicker. It took us almost 5 hours! I was just copying and pasting commands and outputs. On the unix history command recalls everything. I don't recall any, HAHAHA! So what is your process? submitted by /u/Oxffff0000 to r/devops [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.44View on Reddit

Most teams I talked to don't have a CI/CD pipeline checking cost or governance policy before terraform apply, they just run it locally. That's what guard is for: cloudcosttree guard -- terraform apply. To be clear, this isn't a simulation, it's your real terraform apply. guard never runs one on its own, it only wraps the exact command you were already going to run, checks the plan against your policies, then applies that same saved plan, so there's no gap between what got checked and what got deployed. Default behavior is warn-only, it prints violations but still applies. --block opts into actually stopping the apply on a real violation. A false positive blocking a real deploy is worse than one showing up in a report, so blocking is never the default. submitted by /u/Independent-Ease-609 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditAZUREimportance 0.44View on Reddit

I'm still seeing the old GPT pricing in Foundry Monitoring and in my Azure cost management for Standard Global as of today, does anyone know when they will fix this? It's been 12 days now... Support isn't helpful and doesn't have any clue it seems. submitted by /u/googleaddreddit to r/AZURE [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditawsimportance 0.44View on Reddit

Context: I'm running a public API on API Gateway + Lambda, with DynamoDB behind it. Endpoints have very different backend costs, cheap reads vs. a couple of routes that kick off heavier aggregation work. Currently using API Gateway usage plans with a single throttle limit per API key, applied flat across all routes. The problem: usage plans throttle by requests/second regardless of which route is hit, so a client hammering cheap GETs eats the same budget as one calling the expensive routes, and there's no way (as far as I can find) to weight individual routes differently within a single usage plan without splitting them into separate API Gateway stages/plans per cost tier, which gets awkward to manage as the number of "cost classes" grows. What I've looked at so far: Per-stage/per-plan splitting: works, but means maintaining N usage plans and N sets of API keys per client if a client needs access to routes at more than one cost tier. Custom Lambda authorizer + DynamoDB counter: doing weighted token-bucket logic myself (consume different token amounts per route, check/decrement atomically via DynamoDB conditional writes), seems doable but adds a DynamoDB read/write on every request just for the rate-limit check, plus I'd be reimplementing throttling that API Gateway mostly already does for free. Briefly looked at whether Lambda reserved/provisioned concurrency per function could act as an implicit cost-based limiter (route the expensive endpoint through its own function with tighter concurrency), but that limits total throughput, not per-client fairness. Has anyone actually shipped weighted/cost-based rate limiting on top of API Gateway usage plans, or does everyone end up rolling their own with a Lambda authorizer + DynamoDB/ElastiCache counter once costs diverge enough between routes? And if you rolled your own, did you keep API Gateway's built-in throttling as a coarse backstop on top of it, or drop it entirely in favor of the custom logic? submitted by /u/ClickOk5811 to r/aws [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.44View on Reddit

I need some help figuring out if this is a highly targeted scam or a legitimate invoice. I recently received an email containing a PDF invoice for an "Unreturned Advance Exchange Fee". The sender email is ⁠ [email protected] ⁠. Here is what is really tripping me up and making me second-guess: The details of the item on the invoice (a Surface Laptop) and the serial number of the returned item is correct. My personal and organization details listed in the "Bill To" and "Ship To" sections are 100% correct. The invoice claims to be from "MICROSOFT PTY LIMITED" based in North Sydney, Australia, and the total amount is listed in AUD. **The Red Flag:** The "Remit to Bank" section instructs me to send payment to "BANK OF AMERICA". I am based in the Asia Pacific region, so seeing Bank of America as the payment destination feels incredibly suspicious, even though Microsoft is a US-based company. Has anyone else dealt with this before? Is it normal for Microsoft's APAC/Australian branches to use Bank of America for direct wire transfers, or is this just a very sophisticated, highly personalized spoofing attempt using an email address like **⁠ [email protected] **⁠? Customer Success Manager from MS hasn’t replied in a week where I asked if this is legit and I should reply to it. Also btw device was returned within directed time frame. Any advice would be greatly appreciated! submitted by /u/laleric to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.44View on Reddit

terraform plan validates syntax and state, it has no opinion on whether the instance size you picked is actually right for the workload. Ran a single-resource RDS example through CloudCostTree: db.t3.large, declared as-is, comes out to $107.28/month, Cost Score F. Simulating one size down, db.t3.medium, same file, no architecture change: $57.64/month. A real $49.64/mo drop, 46.3%. Nothing about that shows up in a terraform plan diff, you only see it once you actually price the resource. submitted by /u/Independent-Ease-609 to r/FinOps [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditFinOpsimportance 0.42View on Reddit

Hi everyone, so we're looking to better manage our AI spend across multiple dev teams mainly on how to deal with the fluctuating and unpredictable cost of it all. We've already set team / project specific API keys so we can track spend by project, so now I'm just looking for ways to manage the cost itself as a whole. I looked into LLM routers / gateways, mainly from seeing the Ramp Router announcement and it looked interesting to me from a cost cutting perspective. But I have 0 experience in using LLM routers so would love to hear from you guys. I'm open to other suggestions too of course, thanks! submitted by /u/Dangerous_End8856 [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditgooglecloudimportance 0.39View on Reddit

I only have two VM's. I can see in Billing -> Reports I am getting charged $2/day for Persistent Disk. But I can't tell which VM is accounting for most of that charge. submitted by /u/imitation_squash_pro to r/googlecloud [link] [comments]

1 comment
u/olalof

Add labels to the disks and filter by labels in the billing data.

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.37View on Reddit

I know this is obsessive. That is the point. For time, my definition of a clean Windows installation was simple: Create an official Microsoft installation USB. Boot from it. Delete every partition. Install Windows. Sign in with a Microsoft account. Install Windows and Microsoft Store updates. Remove the applications I did not need. Install the drivers from my laptop manufacturer. Apply all my personal settings. Windows was clean, updated, and ready. Or at least, that was what I thought. The standard clean installation was not clean enough The first problem was driver installation. As soon as Windows connected to the internet, Windows Update started installing drivers automatically. After that, I installed the official driver packages from my laptop manufacturer’s website. This meant Windows could install one driver first, and then the manufacturer’s installer could install another version over it. Sometimes everything worked, but the process did not feel controlled. Windows Update also tried to install many updates, drivers, and Microsoft Store components simultaneously. This occasionally produced failed updates, retries, and unnecessary confusion. There was another issue: removing applications after they had already been installed did not feel truly clean. The applications had already been registered, initialized, and possibly updated. Uninstalling them afterward was not the same as preventing them from being installed in the first place. Improvement 1: Complete the first boot without internet My next method was: Install Windows → Complete OOBE using a local account → Enter the desktop without internet → Install the official laptop drivers from USB → Restart → Connect to Wi-Fi → Install Windows and Store updates → Remove unwanted applications → Apply my settings This was much better. Windows Update could not install random drivers before I installed the correct manufacturer drivers. It felt almost perfect. But it still was not perfect. When I finally connected to the internet, Windows still had to process a large cumulative update, smaller updates, driver checks, security updates, and Store updates at the same time. There were still occasional update errors and retries. Improvement 2: Install the large cumulative update offline I decided to integrate the latest large cumulative update into Windows before connecting the computer to the internet. The idea was simple: Install the large update offline → Boot Windows → Connect to the internet later → Download only the smaller remaining updates This significantly reduced the work Windows Update had to perform after the first boot. Instead of downloading and processing a massive cumulative update alongside everything else, Windows only needed the smaller updates released afterward. The process became faster, more predictable, and less likely to produce errors. But I still had two problems: Unwanted applications were being installed before I removed them. Manufacturer driver packages contained drivers for hardware my laptop did not actually have. Improvement 3: Extract and select the exact drivers Many laptop manufacturers provide one driver package that supports several possible hardware configurations. For example, a Wi-Fi driver package may include: Intel drivers Realtek drivers Qualcomm drivers MediaTek drivers My laptop only uses one of them, but the complete package may copy or stage drivers for several supported configurations. That did not feel precise enough. I started extracting the manufacturer’s .exe driver packages and examining the files inside them. I identified the exact hardware installed in my machine and selected only the appropriate drivers. However, I also learned that not every driver should be treated the same way. Some drivers are primarily standard INF-based packages and work well with offline DISM injection. Other packages are more complex and may include: Multiple dependent drivers Services Registry configuration Software components Firmware utilities Control panels Microsoft Store or UWP companion applications Custom installation logic Graphics, audio, and some Intel platform or firmware packages can fall into this category. For those packages, blindly extracting every INF file and injecting everything is not necessarily correct. I therefore inspected each driver package and divided them into two groups: Safe standard drivers → Inject offline with DISM Complex software-assisted drivers → Install after the first desktop boot using the official installer This gave me control without breaking the functionality supplied by the manufacturer. Improvement 4: Inject drivers before Windows boots Even when the correct driver was installed, installing it after reaching the desktop could temporarily restart or reinitialize the related device. For example, installing a network driver can cause the adapter to disappear and reappear. That is completely normal, but I wanted the hardware to be ready from the first real Windows boot. I therefore started servicing install.wim offline with DISM. My early method was: Copy install.wim from the official Microsoft ISO → Mount the WIM → Inject the safe drivers → Inject the cumulative update → Modify the image → Commit and unmount it → Use the modified image for installation DISM is already part of Windows, so the entire servicing stage could be performed using Microsoft’s own deployment tools. At this point, the installation already had the correct core drivers and the largest Windows update before it ever reached the desktop. Improvement 5: Prevent unwanted Store applications from being provisioned Microsoft Store applications are commonly provisioned in the Windows image. Provisioned applications are prepared so that Windows can register them when a new user account is created. Instead of allowing these applications to register and then uninstalling them afterward, I removed the unwanted provisioned packages from the offline image using DISM. First, I listed the provisioned packages: dism /Image:W:\ /Get-ProvisionedAppxPackages Then, for each package I did not want: dism /Image:W:\ /Remove-ProvisionedAppxPackage /PackageName:<exact-package-name> This meant the applications were removed before my user profile was created. I did not remove essential components such as Microsoft Store or Desktop App Installer. I only removed applications I had already decided I would never use. This felt much cleaner than uninstalling them after the first login. Improvement 6: Prevent OneDrive Setup from starting for the new user OneDrive was a separate case. It was not simply a provisioned Store application in the same way as the other packages. Windows contained a startup entry that launched OneDrive Setup when the user profile was created. I loaded the offline default-user registry hive and removed the OneDrive Setup startup entry. That stopped OneDrive Setup from automatically running when the first user account was created. Again, the objective was not to install something, allow it to initialize, and then remove it. The objective was to prevent the unwanted setup process from starting at all. Improvement 7: Apply my settings before the first login I then realized that many settings could also be applied offline. Instead of entering the desktop and manually changing everything, I loaded the offline registry hives and configured settings such as: Dark application mode Dark Windows interface mode Windows Spotlight policies News and Interests Delivery Optimization Fast Startup OneDrive startup behavior Other system and default-user preferences The settings were therefore present when Windows created the user profile. The system started in the state I wanted instead of starting with default settings and being changed afterward. Everything was now: Official Controlled Repeatable Performed with built-in Windows deployment tools Completed before Windows created its first normal user session But I still was not satisfied. The normal USB installer still hid too much of the process Even with a customized install.wim, the normal Windows Setup interface was still performing many operations automatically. It created partitions, applied the image, configured the boot files, and prepared recovery behind the scenes. That is convenient for normal users. For my perfection-obsessed installation, however, I wanted to know and control exactly what was happening. So I stopped using the normal graphical installation process. I moved the entire deployment into WinRE. The final method: Build Windows manually from WinRE WinRE is a small recovery environment that boots into RAM and provides access to tools such as DiskPart, DISM, BCDBoot, Registry Editor, and ReAgentC. I placed the official Microsoft install.wim, updates, drivers, answer file, and recovery image on a separate drive. Then I booted into WinRE and manually built the complete Windows disk. Step 1: Create the GPT partition structure manually I selected the correct target disk, deleted the existing partition structure, and manually created: EFI System Partition — 300 MB FAT32 Microsoft Reserved Partition — 16 MB Windows partition Windows Recovery partition — 2048 MB This gave me full control over the partition sizes, order, and purpose. The EFI, Windows, and Recovery partitions were assigned temporary drive letters during deployment. Step 2: Apply the official Windows image Instead of running Windows Setup, I applied the Microsoft image directly: dism /Apply-Image /ImageFile:C:\install.wim /Index:1 /ApplyDir:W:\ At this stage, W: contained a fresh Windows installation that had never booted. Step 3: Service the offline installation While Windows was still completely offline, I: Injected the compatible drivers Installed the latest large cumulative update Removed unwanted provisioned applications Modified the default-user and system registry hives Disabled the OneDrive Setup startup entry Applied my preferred Windows settings Added the unattended configuration Cleaned the component store For the final component cleanup, I used: dism /Image:W:\ /Cleanup-Image /StartComponentCleanup /ResetBase I only used /ResetBase after finalizing the update state because it prevents the integrated updates from being uninstalled later. Step 4: Configure OOBE officially through Panther A recent Windows 10 update complicated the local-account path during OOBE, and Windows 11 also strongly encourages online account creation. There are command-line workarounds that restart or interrupt OOBE, but that did not feel clean to me. Instead, I used a Windows answer file in the Panther directory: W:\Windows\Panther\unattend.xml The answer file configured items such as: Language Keyboard layouts Time zone Local user account OOBE behavior Automatic initial login This is part of Windows Setup’s own unattended-deployment system. It did not require third-party bypass tools or interrupting OOBE with an improvised restart. Any plaintext password contained in the answer file should be removed after Setup is complete. Windows 11 builds can behave differently, so an answer file should be validated against the specific build being deployed. Step 5: Create the UEFI boot files manually I generated the boot environment directly from the applied Windows installation: W:\Windows\System32\bcdboot.exe W:\Windows /s S: /f UEFI This created the UEFI boot files on the EFI System Partition. Step 6: Build and register the recovery environment I created the recovery directory, copied winre.wim, and registered it against the offline Windows installation: md R:\Recovery\WindowsRE copy /y C:\winre.wim R:\Recovery\WindowsRE\winre.wim W:\Windows\System32\reagentc.exe /setreimage /path R:\Recovery\WindowsRE /target W:\Windows I then assigned the proper recovery-partition GPT type and attributes and removed its temporary drive letter. The result was a complete disk containing: A fresh UEFI boot partition Windows A properly configured recovery environment No previous user activity No Audit Mode session No Sysprep generalization cycle The controlled first boot The offline deployment was only the first stage. For the first real boot, I still kept the laptop disconnected from the network. My sequence was: Boot Windows without internet → Complete the local OOBE process → Enter the desktop → Install the complex official driver packages → Restart → Pause Windows Update temporarily → Connect to Wi-Fi → Allow required OEM and Store components to install → Run Windows Update → Update Microsoft Store applications → Apply the remaining personal settings → Install required DirectX and Visual C++ runtimes → Restart again The complex driver packages included components that were better installed through their official installers, such as graphics, audio, and certain Intel platform packages. After connecting to the internet, official companion applications such as Dolby or Intel software could install normally. Because the large cumulative update had already been integrated offline, Windows Update only had a relatively small amount of remaining work. Because unnecessary provisioned applications had already been removed, Microsoft Store also had fewer applications to register and update. At this point, I had a complete Windows installation with: Official Microsoft Windows files Official manufacturer drivers Official NVIDIA drivers Official Microsoft updates Official Microsoft Store components My complete configuration No random third-party customization utility No unnecessary driver families No unwanted provisioned applications A working EFI partition A working WinRE partition This was finally the Windows state I wanted. But there was one final problem. Reproducing all of this again would take hours. Preserving the perfect state Once I started using Windows normally, installing random applications, testing software, and modifying files, the installation would no longer be in that carefully prepared state. I wanted to preserve it before normal use. Windows includes the older system-image backup concept, and many third-party disk-cloning tools also exist. However, I had completed the entire deployment using Windows-native tools. I did not want the final backup stage to depend on an unrelated third-party cloning application. So I returned to WinRE and used DISM again. WIM capture First, I captured the Windows partition into a WIM file: dism /Capture-Image /ImageFile:C:\Final-Windows.wim /CaptureDir:W:\ /Name:"Final Ready-to-Use Windows" /Compress:max /CheckIntegrity /Verify Because the capture was performed from WinRE, the Windows installation was offline. The WIM represented the files on the Windows partition at the exact point when I shut the system down. Restoring it later would return the Windows partition to that state. However, a WIM only captures the selected partition. It does not automatically preserve: The EFI partition The Microsoft Reserved partition The recovery partition The complete disk layout To restore a WIM to a completely empty disk, I would still need to recreate the partitions, apply the WIM, rebuild the boot files, and configure recovery again. That was not perfect enough for my objective. FFU: The complete-disk image DISM also supports Full Flash Update images. Instead of capturing only the Windows partition, an FFU captures the physical disk layout. From WinRE, I used: dism /Capture-FFU /ImageFile:C:\Final-Windows.ffu /CaptureDrive:\\.\PhysicalDrive1 /Name:"Final Perfect Windows" The FFU contains the complete disk: EFI partition MSR partition Windows partition Recovery partition Partition order Boot files Windows files Drivers Updates Applications Settings Recovery configuration Now I can boot into WinRE, erase the target disk, and apply the FFU. It recreates the complete disk structure and restores the files to the captured state. Instead of repeating hours of partitioning, servicing, driver selection, updating, and configuration, I can restore the complete prepared Windows environment in a small fraction of the time. It is effectively like returning the machine to the exact day when I finished preparing it. The WIM remains useful as a flexible Windows-partition backup. The FFU is the complete bare-metal recovery image. The final result My complete workflow became: Official Microsoft ISO → Manual GPT partitioning in WinRE → Direct DISM image application → Offline cumulative update integration → Selective offline driver injection → Offline provisioned-application removal → Offline registry customization → Official unattended OOBE configuration → Manual UEFI boot creation → Manual WinRE registration → Component-store cleanup → Controlled offline first boot → Official complex driver installation → Controlled internet connection → Remaining Windows and Store updates → Final runtimes and settings → Offline WIM capture → Complete-disk FFU capture Is this necessary for most people? Absolutely not. The normal Microsoft installer is sufficient for almost everyone. But for someone who is obsessed with understanding, controlling, and perfecting every stage of a Windows installation, this is the closest I have reached to an absolutely clean, official, and reproducible system. The best part is that I can now use Windows freely, install experimental software, and potentially break things without worrying about repeating the entire process. My final FFU image preserves the complete perfect state. I only need to restore it, and I am back where I started. Important notes Always verify the physical-disk number before using DiskPart or FFU commands. Selecting the wrong disk can destroy data. Test the FFU restoration before treating it as your only backup. FFU images are less flexible than WIM images and are best suited to the same machine, disk layout, or compatible target storage. The target disk generally needs enough capacity for the captured layout. Some driver installers should not be replaced with blind INF injection. Windows 10 and Windows 11 do not always use identical package names, updates, or unattended settings. Remove answer files containing passwords after Windows Setup completes. Keep personal files backed up separately. A system image is not a substitute for a separate data backup. I am interested in hearing how deployment specialists would improve this workflow, and whether anyone else has taken the Microsoft-native WinRE, DISM, WIM, and FFU approach this far. submitted by /u/kaidocodm to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X
redditsysadminimportance 0.37View on Reddit

Renewal season is hitting a lot of people and I keep seeing the same pattern: the quote gets escalated before anyone has actually measured what they run. RVTools is free, read-only and takes about twenty minutes with a service account. Here is the checklist I would work through before replying to anyone. Count billed cores, not sockets. The metric moved from per-socket to per physical core. vHost gives you "# CPU" (sockets) and "Cores per CPU". Multiply, sum per site. Then apply two floors, in this order: The historical 16-core-per-socket minimum. An 8-core CPU bills as 16. The 72-core-per-product-per-order-line minimum, which went up from 16 in April 2025. The second one is the one people miss, and it lands on small sites rather than big ones. 16-host cluster, dual 32-core CPUs = 1,024 cores. Neither floor touches it. 3-host edge site, one 16-core CPU per host = 48 real cores, bills at 72. A 50% overpay. 2-host site, dual 8-core CPUs = 32 real cores, floored to 64 by the socket rule, then to 72 by the order-line rule. A 125% overpay on a box in a cupboard. Worth confirming with your reseller whether they can consolidate sites onto one order line, because that changes the answer a lot. Get it in writing. Find what you are licensing for no reason. vInfo, filter Powerstate = poweredOff, sum "In Use MiB". Powered-off VMs still occupy storage and still sit in your capacity planning. vSnapshot, anything with a date older than about 30 days. Those are a production risk as well as reclaimable space. vHealth has a built-in check for possible orphaned VMDKs and zombie files. Provisioned is not consumed, and that is the number that decides your business case. Sum CPUs in vInfo for powered-on VMs, divide by total physical cores from vHost. That is your real vCPU:pCore ratio. Do it against physical cores, not hyperthreaded logical CPUs, because licensing is physical. Then compare three storage numbers that usually get treated as one: provisioned VMDK capacity (vDisk), datastore in-use (vInfo), and guest filesystem consumed (vPartition). The gap between the first and the third is often most of your "capacity". This matters beyond the renewal. Cloud block storage bills on allocated volume size, not what the guest is using, so a thin 2TB VMDK holding 200GB becomes a 2TB bill on day one unless you shrink it first. Sizing any target on provisioned figures will make a migration look far more expensive than it actually is. Gotchas worth knowing before you start: Column names drift between RVTools versions, and older exports say MB where newer ones say MiB. Print your columns before writing any filters. RVTools is a configuration snapshot, not performance history. It cannot tell you utilisation over time. You still want around 30 days of vCenter stats and a P95 per VM before you size anything. Check vDisk for independent/persistent disk mode and RDMs. Both break most snapshot-based migration tooling. Check vMemory for ballooned and swapped. That tells you where you are already past comfortable overcommit. RVTools cannot see inside the guest, so it will not tell you where SQL Server is installed. On some estates the Windows and SQL per-core licensing delta is bigger than the hypervisor saving, and it follows you to whatever you migrate to. Use a dedicated read-only account, not an admin one, and treat the export as sensitive. It is a complete map of your estate. One framing thing. Moving to a managed vSphere offering on a hyperscaler is not an exit, it is a change of landlord. It may still be the right call if you are up against a datacentre lease deadline, but it is worth saying out loud in the writeup before someone else does. Also worth pricing honestly: "renew a smaller, cleaned-up footprint" is a legitimate option, and doing the cleanup makes every other option cheaper too. Curious what ratios other people are finding when they actually pull the numbers. The provisioned-to-consumed gap seems to be the one that surprises people most. submitted by /u/RulezZzOr to r/sysadmin [link] [comments]

Repurpose (generate each channel independently)
Discord
LinkedIn
X