Volume Without Value – Part Two: Which Of Your Security Data Is Earning Its Place?
Cutting security data to save money is easy to do badly. Here is a method for working out what is actually earning its place, and where everything else should live instead.
Part one of this series argued that the growth in security data volume is being driven by ordinary organisational change rather than by any increase in risk, and that no amount of negotiation fixes an architectural problem. The obvious follow up question is which data you can move or stop paying premium rates for without opening a hole in your coverage.
The honest answer is that most organisations do not know, because the work required to find out has never been prioritised. It is not difficult work. It is simply nobody’s job.
Start With What You Cannot Currently See
Before any decision about tiering or filtering, you need three pieces of information that most teams cannot produce on request.
- A mapping of every active detection rule to the data sources it depends on. Not the sources it was written against originally, the sources it queries today.
- A record of which sources analysts have actually queried during investigations over a meaningful period. Ninety days is usually enough to see the pattern.
- A list of every connected source with its current volume, its retention setting and the reason it was connected in the first place.
The third item is the one that produces the surprises. In most estates there are sources connected years ago for a project that has since concluded, sources duplicated because two teams onboarded the same platform, and sources nobody can now account for at all.
The Four Categories
Once you have that picture, every source and in many cases every field within a source falls into one of four groups.
Category | What it is | Where it belongs |
Detection | Data that actively fires rules or feeds correlation logic | The searchable tier. This is what you are paying for and it is worth it |
Investigation | Data analysts reach for during an incident but which triggers nothing on its own | Accessible storage with acceptable query latency. Rarely needs premium tier |
Obligation | Data retained to satisfy a regulatory, contractual or audit requirement | Low cost durable storage with a documented retrieval path |
Dormant | Data nobody has queried since it was connected | Question whether it should be collected at all |
The distribution varies considerably between organisations, and I would be cautious of anyone quoting you a standard ratio. What is consistent is that the detection category is smaller than people expect, and the dormant category is larger.
Where The Data Should Go Instead
Moving data out of the premium tier only helps if it lands somewhere it remains usable. Cheap storage that renders data unqueryable in practice has not solved the problem, it has converted a cost problem into a capability problem and deferred the consequences to your next incident.
Two things determine whether the destination is any good.
The first is the format the data lands in. If it is written in a proprietary or tool specific structure, you have created a dependency that will outlast whichever platform prompted it. Open schemas exist for exactly this reason. The Open Cybersecurity Schema Framework has gathered meaningful support across the industry, and normalising to an open standard on the way in means the data remains readable regardless of what sits on top of it later.
The second is who owns the storage. Data held in your own cloud account, in a security data lake you control, behaves differently at renewal time from data held inside a vendor platform. That distinction matters commercially as well as technically, and part three of this series looks at why.
The Compliance Objection
The most frequent objection to any of this is that regulatory obligations require the data to be retained, so it cannot be touched.
Retention obligations generally specify how long data must be kept and how reliably it must be produced when requested. They very rarely specify that it must be instantly searchable in a premium analytics platform for the duration. Those are different requirements and conflating them is expensive.
What matters is that the retrieval path is documented, tested and defensible. An obligation satisfied by low cost durable storage with a proven restore process is satisfied. This is a point worth confirming with your own compliance function rather than taking on trust from a blog post, including mine, because the specifics vary by sector and by regulator.
Two Ways This Goes Wrong
The first failure is cutting a source that was quietly feeding several correlation rules. This is why the rule to source mapping comes before any decision rather than after it. Detection logic accumulates over years and the dependencies are rarely documented anywhere a procurement exercise would find them.
The second is subtler and more common. Teams reduce volume by dropping fields rather than sources, on the reasonable assumption that a partial record is better than none. Sometimes it is. But field level filtering breaks detections in ways that are much harder to spot than source level removal, because the rule still runs, still returns results and simply misses a subset of cases. A detection that fails loudly is a problem. A detection that fails quietly is a much worse one.
If you take one operational principle from this piece, make it this: change one thing at a time, record what you changed, and validate the affected detections before moving on.
What Changes When This Is Done Properly
The immediate result is a lower bill, which is what gets the project approved. The more durable result is that the security function regains control of a decision it had effectively outsourced to a licensing model.
Once data is being routed deliberately rather than by default, connecting a new source becomes an architectural decision with a known cost rather than a budget event. Teams stop declining coverage for financial reasons, which is the outcome that actually matters.
In part three we look at what all of this does to the renewal conversation, and how to measure whether the work delivered what it promised.
HOOP Cyber helps security teams audit their data estate, design the routing and normalisation that sits behind it, and rebuild security operations around a data architecture they control. If you cannot currently produce the three pieces of information at the top of this article, that is where we would start. Get in touch via .