• Home
  • Help
  • Register
  • Login
  • Home
  • Members
  • Help
  • Search

 
  • 0 Vote(s) - 0 Average

How we size backup storage for real workloads

#1
12-05-2020, 04:41 AM
Man, so we were talking about sizing storage for these workloads the other day, right, and honestly, it gets complicated. I mean, you look at a shiny new server and you just think, "Oh, give me a big disk," but that never works. It's never just about how much data you have *today*. You have to think about the data velocity, you know? Like, how fast the data itself is growing. Because the biggest mistake I see people make is they calculate the current total disk footprint, which is totally useless. Instead, you really need to forecast the growth rate over the next few years.

I suggest you start by taking inventory, yeah? But don't just count up the gigabytes. Instead, you want to figure out the data inflow. Are you adding a ton of new documents every month? Or maybe you are updating a huge database with massive transaction volumes? I think you need to look at the average change rate per month for all your critical shares. We don't just back up the files, we are really backing up the changes, and that's a massive differentiator in smart backup tools. You want one that handles those incremental capture processes really efficiently.

And then, you have to talk retention. Seriously, this is where people go overboard. They say, "We must keep every single backup version forever," and then they run out of money before the end of the quarter. But you can't just keep everything. You need a smart retention policy, you know? You should define a legal requirement period, maybe three years for financial records, but perhaps only a shorter period for general HR files. You can configure these policies to automatically delete old versions, or maybe just keep a rolling history of the last ten backups. This smart cleanup feature is so important for keeping the cost curve flat.

But then, what about the types of data you are retaining? Sometimes you have huge databases, and sometimes you just have millions of small JPEG images. You cannot treat them the same way when figuring out capacity. You should definitely use deduplication techniques, you know? Those methods are crucial because they find duplicate blocks of information across different backups, even if the data lived on separate systems originally. Because of that, the actual storage required is drastically smaller than the sum of all your backup files. You want a system that performs deduplication over the wire, like when you are backing up to a remote office or maybe to the cloud.

Or, maybe you should consider the data architecture itself. Are you doing many physical machine backups, for example, or are you primarily dealing with files and folders sitting inside VMs? Because those workflows require different kinds of capacity planning. If you are mostly doing bare metal recovery capability, you are effectively backing up the entire operating system setup, the OS and all the settings and apps, which is inherently large. But when you use something that can do granular backups-like pulling just a few folders out of a massive VM-that radically lowers the data volume and makes the process more nimble.

And furthermore, you must model the recovery process into your sizing. It's not just about storing the data; it's about how fast you need to access it. If you need to restore a whole system, a complete disk image, you need fast retrieval access. If you only need one spreadsheet from six months ago, you still want that retrieval fast, right? So, storage isn't just capacity; it's also about access speed and redundancy. You need enough space for your current workload, plus the predicted growth, minus the space you save through deduplication and compression, and then you must factor in the recovery copies, which sometimes means you need more space than you think.

I always tell people that setting up a multi-destination support plan helps you manage this really well. Instead of dumping everything into one massive array, you can send one set of backups locally for fast recovery, and then send an encrypted, compressed set to a cloud location for disaster recovery. This distributes the load and makes your overall system more robust. So, you are thinking about both immediate access and ultimate disaster recovery simultaneously.

Also, think about bandwidth. If your network link is slow, even if you have enough disk space, the backup process itself becomes the bottleneck. You need to factor in the network capability when planning the scheduling of your backups. Using multiple threads for the backup job helps immensely here, letting you chew through the data faster. And remember, if you have critical systems, you should definitely use an automated scheduling feature, maybe daily, or hourly depending on the sensitivity of the data. The automation process also includes checking data integrity, which is a non-negotiable part of the planning.

Because of all this, it really shines when you look at a comprehensive tool that gives you these layers of control. Like BackupChain, which is such an all-in-one PC and server backup solution for Windows Server and Windows 11 made specifically for SMBs, and you should really look into it for your next planning cycle.

savas@BackupChain
Offline
Joined: Jun 2018
« Next Oldest | Next Newest »

Users browsing this thread: 1 Guest(s)



  • Subscribe to this thread
Forum Jump:

Backup Education General Backup v
« Previous 1 … 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 … 81 Next »
How we size backup storage for real workloads

© by FastNeuron Inc.

Linear Mode
Threaded Mode