
For years, security and IT teams have made the same uncomfortable trade with network telemetry: keep everything and pay for storage and indexing you cannot justify or drop most of it and lose the visibility you need when an investigation reaches back weeks or months. Full fidelity was expensive. Affordability meant gaps.
At Cisco Live 2026, Splunk moved to break that trade with the Machine Data Lake, a storage foundation, currently in Alpha, where machine data is automatically cataloged, enriched, and governed at scale without requiring immediate indexing. The pitch is direct: preserve full-fidelity data, cut the indexing costs that made retention painful. Splunk cites a global insurance company that modeled roughly 50% lower spend while keeping full-fidelity retention.
This is a meaningful shift in the economics of keeping data. But cheaper storage raises a second question that matters just as much: if you are going to retain everything, is what you are retaining actually worth keeping? For network telemetry, that is where the real work happens before the data ever reaches the lake.
Cheaper Storage Does Not Fix Bad Data
A data lake changes what you can afford to store. It does not change what the data is worth. Raw network flow, retained at full fidelity in a low-cost lake, is still raw network flow: source and destination IPs, ports, byte counts, timestamps. Retaining more of it, more cheaply, gives you a larger pile of thin records.
When an investigation reaches back into that retained data, the value depends entirely on what each record contains. A flow record that shows only that 10.1.4.22 talked to an external IP for four minutes is nearly useless six weeks later, when no one remembers who was behind that address or what application was involved. A record that already carries the user identity, the application, the threat intelligence context, and the geographic origin is an audit-ready artifact whenever you return to it.
A cheaper lake makes it affordable to retain everything. It does nothing to make what you retain useful. The value of retained telemetry is set before it lands, by how much context each record carries.
Enrich and Reduce Before the Lake, Not After
There are two ways to get network telemetry into a storage foundation like the Machine Data Lake. You can dump raw flow in and try to enrich it later at query time, or you can enrich and reduce it on the way in, so that what lands is already complete and already free of redundant bulk.
The second approach is the one that pays off, for two reasons. First, query-time enrichment of raw flow is slow and fragile: the identity mappings, threat feeds, and application context that were current when the traffic occurred may be gone or changed by the time you query months later. Enrichment must happen when the flow is live to be accurate. Second, raw flow carries enormous redundant volume, especially the repetitive machine-to-machine conversations that dominate modern networks. Reducing that volume before storage means the lake holds meaningful records, not noise.
NetFlow Optimizer (NFO) does exactly this, in the pipeline, before delivery. It parses the binary NetFlow, IPFIX, sFlow, and J-Flow exported by routers, switches, and firewalls, normalizes it to a common information model, enriches every record with user identity, application, threat intelligence, and geographic context, and reduces volume by 80 to 90% through aggregation. What lands in the lake, or in Splunk, Sentinel, or another destination, is already enriched, already structured, and already free of the redundant bulk that would otherwise inflate storage without adding insight.
| Raw flow into the lake | NFO-processed flow into the lake |
| Thin records: IPs, ports, bytes | Enriched: identity, app, threat intel, geo |
| Enrich later, if the context still exists | Enriched live, when the context was accurate |
| Full redundant volume stored | 80 to 90% reduced via aggregation |
| A large pile of low-value data | Audit-ready records worth retaining |
The Two Reductions Compound

The Machine Data Lake reduces the cost per unit of stored data. NFO reduces the number of units worth storing and raises the value of each one. These are complementary, and they compound.
Consider the sequence. NFO aggregates raw flow, cutting volume by 80 to 90% before anything is stored. The enriched records that remain then land in a storage foundation that is far cheaper per unit than traditional indexed retention. The organization ends up retaining full-fidelity, fully-enriched network telemetry, for the long windows that investigations and compliance require, at a fraction of what either mechanism would achieve alone. This is also why full-fidelity network telemetry becomes viable to keep in Splunk in the first place, rather than being dropped to control cost.
The same logic extends across the Cisco Data Fabric, the Splunk-powered layer Cisco positions as the system of record for the enterprise. Federated search across that fabric is only as useful as the quality of what it searches. Enriched, structured network telemetry is a first-class citizen of that data layer. Raw binary NetFlow is not.
The Bottom Line
The Machine Data Lake changes what full-fidelity retention costs. It does not change what raw network telemetry is worth. Retaining thin, unenriched flow records more cheaply still leaves you with thin, unenriched records when an investigation reaches back for them. The value of retained telemetry is decided before it lands, by how much context each record carries and how much redundant volume was stripped out first.
NFO does that work in the pipeline: enrich live, reduce by 80 to 90%, deliver records that are worth keeping. Cheaper storage and better data are not competing strategies. Used together, they are how full-fidelity network visibility finally becomes something you can afford to keep.
Rethinking what network telemetry costs you to retain? Start a free 60-day trial of NetFlow Optimizer or schedule a technical demo to see enriched, volume-reduced flow telemetry in action.
Start Free Trial | Schedule a Demo | Reduce SIEM Ingest | NFO Documentation
