
The Firehose Is Not Intelligence
Originally published on Medium.
Intelligence platforms have spent years competing to collect more sources and process more content. As synthetic and derivative material grows, that architecture increasingly risks confusing repetition with evidence. The next generation of systems should be designed around reconstructing events, preserving provenance and retrieving information when it can materially change the assessment.
Consider an analyst is reviewing reports about a disruption at a large industrial site. In 60 minutes time, they are expected to advise whether the disruption is serious enough for the organisation to warn customers and begin looking for alternative suppliers.
At first, there is not much information. A photograph shows several vehicles outside an entrance. One local source says workers have been sent home, while another claims production has stopped. Within an hour, dozens of accounts are repeating the story. By the end of the day, thousands of posts, articles, reactions, screenshots and translations describe what appears to be a shutdown.
A system designed to gather information is doing exactly what it was built to do. It downloads the photograph in several resolutions, translates posts, groups similar articles, rates accounts and extracts the details that appear to matter. As the volume grows, dashboards show increasing activity. To an outsider, it looks as though intelligence is being created in real time.
But when the analyst traces the reporting back to its origin, almost everything points to the same photograph and the same unverified post. What looked like thousands of observations was really one observation spreading through a network.
The amount of content was real. The amount of evidence was not.
Meanwhile, several quieter details are easy to overlook. A local authority has announced a temporary road restriction near the site. Shipping activity associated with the facility has fallen during the previous week. A supplier has changed its delivery schedule, and a maintenance contractor has cancelled several open shifts.
None of these details proves that the facility has closed. None is especially dramatic in isolation. Together, however, they begin to describe a meaningful disruption through several independent lines of evidence.
The analyst's decision depends less on how many items the system has collected than on whether it can show how those items relate to one another.
That is the difference between collecting content and reconstructing a situation. It is also where many intelligence products become trapped between the system they assume they need and the one they may not need to build at all.
More access, but not necessarily more understanding
Much of the intelligence industry still appears to compete on access. Companies advertise the breadth of their coverage, the number of sources they monitor and the quantity of material they can process each day.
The premise is understandable. Important events leave traces, and many of those traces now appear in publicly or commercially available information. If one source provides useful evidence, access to more sources should increase the chance of finding something important right?.
The problem begins when this idea becomes an unquestioned assumption that more sources, more posts and more data will necessarily produce better intelligence. The scale of the open-source ecosystem gives some indication of the scale involved. Despite this abundance, finding the right capability is still difficult. (we've all seen a post about someone making a 'Palantir' clone — they haven't btw.)
The commercial-data landscape creates a similar problem. There are more providers, feeds and specialist datasets than most organisations can reasonably evaluate, integrate or continuously monitor. Some offer genuinely distinctive access. Others package information available elsewhere, sometimes through several layers of intermediaries.
More choice can improve coverage, but it also moves the constraint. Someone still has to establish which source is appropriate, whether a vendor has unique access, whether two datasets share the same upstream origin and whether a new feed provides independent evidence or another version of something already known.
The technology is undoubtedly becoming better at finding fragments. Search is faster, translation is cheaper, computer vision is more accessible, and automation can monitor far more material than a person could examine manually. Yet the analytical role has not disappeared. The US Intelligence Community's OSINT strategy describes the discipline as increasingly important and growing in demand, while identifying the development of the next generation of OSINT professionals and tradecraft as a strategic priority.
This is not as contradictory as it first appears. Better tools reduce the cost of finding material, but they do not remove the need to determine what that material represents. By producing more results, alerts and candidate sources, they can move the difficult work further downstream.
There is also an economic problem. Every permanent source introduces another dependency. Formats change, authentication expires, commercial terms are renegotiated, and platforms remove interfaces that had previously been stable. Everything collected must then be moved, normalised, translated, deduplicated, indexed, secured and retained.
Initially, this machinery is treated as a way of reaching an analytical goal. Over time, maintaining the machinery can become the goal itself.
Generative technology makes that architecture even harder to justify. One rumour can now produce thousands of articles, summaries, translations and reactions. The outputs may look different, but they need not contain any new observation. Automated accounts can respond to automated articles that were themselves generated from other automated summaries.
A platform may therefore become more active, more expensive and more confident while learning almost nothing new.
The problem is not that intelligence vendors have never heard of deduplication, clustering or provenance. Many already offer versions of these capabilities. The deeper issue is that breadth of collection is often treated as the durable asset, while reconstruction remains a downstream feature applied after the data has already been acquired.
As the cost of producing content falls, that order becomes increasingly questionable.
The event is outside any single record
The industrial disruption in the opening scenario is a real-world event, but the argument applies equally to digitally native events. A cyber intrusion, an influence operation or a failure in a software service may occur largely through digital systems. Yet, the event still exists outside any single log entry, post, alert or database record.
The world does not arrive as a completed record.
A photograph captures one moment. A road notice records a decision by a local authority. A shipping dataset reflects a change in movement. A cancelled shift shows that someone altered an operational plan. Each source provides a partial view shaped by how the information was created, collected and preserved.
No database contains the shutdown in its complete form. No vendor can sell the entire situation as a finished object. This is why I have come to see intelligence as a reconstruction problem.
In software, we are comfortable working with systems in which the object we care about is not stored directly. An event-sourced application may preserve a sequence of changes rather than a complete current record. A document system may store patches instead of every historical version. A file can be divided into chunks that only become meaningful when assembled according to a manifest.
In each case, the complete object is distributed across smaller pieces, but those pieces contain enough structure for it to be reconstructed.
Intelligence often works in the same way. What matters is rarely an individual post or record. It is the event, relationship, intention or change in the world that the fragment only partially reveals.
A company does not become financially distressed because someone writes that it is distressed. The distress exists independently and may reveal itself through delayed payments, leadership departures, reduced procurement, legal filings, declining activity or unusual communication from management.
No single fragment has to explain everything. The analytical task is to connect the pieces into a defensible and revisable account of what is changing.
That account is what I mean by a narrative. I do not mean choosing the most coherent story or asking a language model to turn a large pile of documents into fluent prose. I mean maintaining a testable model of the actors, events, relationships and uncertainties involved in a situation.
A useful reconstruction places observations in time, distinguishes independent evidence from repetition and remains open to revision when new information appears. It should become less confident when apparent corroboration turns out to have a single origin, and it should show what evidence would cause the current assessment to change.
Once this becomes the goal, the stream of posts begins to look less like the final product and more like an expensive intermediate step. The analyst's role also becomes easier to explain. Analysts are not needed because the search tools are poor. They are needed because finding fragments and understanding events are different kinds of work.
Architecture should follow the question
Changing the unit of intelligence from the post to the event has practical architectural consequences.
A collection-first system asks how it can ingest another source. It measures coverage in feeds, documents, posts and data volume. Its implicit objective is completeness, even though true completeness is impossible.
A reconstruction-first system starts with the question the organisation is trying to answer. It maintains a provisional event model containing the current claims, the evidence supporting them, the relationships between those claims and the uncertainties that still matter.
From there, the operating loop is relatively clear.
The system identifies which unresolved question would most affect the assessment. It then retrieves from sources capable of answering that question. New evidence may support an existing claim, contradict it, introduce a different event or reveal that two apparently independent reports share the same origin. The event model is updated, while provenance links preserve how each conclusion was reached.
In the industrial example, the question may initially be whether operations have stopped. Once several independent indicators support a disruption, another thousand social posts add little. The next meaningful uncertainty may concern duration, cause or impact. Retrieval should then shift towards sources capable of resolving those questions, such as transport records, regulatory notices, supplier communications or observations tied to the facility.
This does not eliminate continuous collection. Some sources are important enough to justify permanent access, particularly when they offer time-sensitive or irreplaceable information. Complete corpora are also necessary when the corpus itself is the subject of analysis, when records must be preserved for legal reasons or when future questions cannot reasonably be anticipated.
Historical data has option value. Evidence that appears unimportant today may become essential after a later event changes the question being asked.
The argument is therefore not that organisations should collect only what fits their current theory. It is that permanent collection should have an explicit analytical, operational, legal or archival purpose rather than being treated as an unquestioned default.
For many other sources, retrieval can be driven by evidence gaps. Instead of maintaining fifty fragile integrations, a team might keep direct access to a smaller set of dependable feeds and query additional sources when they can materially affect the current assessment. Rather than storing every copy of a claim, it might retain the earliest known appearance, a representative sample of derivatives and the relationships between them.
A provenance graph could show how claims, documents and observations are connected, helping analysts distinguish genuine corroboration from repetition. Source prioritisation could consider past reliability, relevance and the likelihood of providing unique evidence, although none of those scores should become a permanent judgement.
Hashes can demonstrate that stored artefacts have not changed, but they cannot prove that the artefacts were accurate or authentic in the first place. Similarly, targeted retrieval can reduce waste, but it can also reinforce an existing theory if the system only looks for evidence capable of confirming it.
A credible design must therefore search deliberately for contradiction. It should record alternative explanations, lower its confidence when expected evidence is absent and periodically query outside the source set suggested by its current model.
The stopping condition should not be that the system has found enough material to tell a smooth story. It should be that further retrieval is unlikely to change the decision at hand, within the time and confidence constraints of the task.
That is a far more useful measure of sufficiency than the number of records processed.
The danger of a convincing reconstruction
A system designed to reconstruct events can still construct the wrong one.
Fragments may be missing. Two apparently independent sources may share an undisclosed origin. Entity-resolution systems may connect different people or organisations with similar names. A plausible sequence of events may appear inevitable only because contradictory evidence was never retrieved.
Language models make this risk more serious. They are extremely good at converting incomplete material into fluent explanations, and fluency can hide uncertainty. A weak conclusion presented smoothly may feel more authoritative than a stronger conclusion presented cautiously.
For that reason, reconstruction cannot simply mean asking a model what happened. A fluent account without visible provenance is another piece of content whose assumptions and origins must be questioned.
Claims should be traceable. Repetition should not be mistaken for independent support. Contradictions should remain visible, and missing information should lower confidence rather than being quietly replaced by plausible guesses.
A useful system should be able to explain not only what it currently believes, but why it believes it, what remains uncertain and what evidence would cause it to change its mind.
Return to the analyst deciding whether to warn customers and look for alternative suppliers.
A weak system reports that the factory has closed because thousands of posts say so. A stronger system reports that operations are likely to have been disrupted, supported by reduced freight activity, a local road restriction and two independent reports from people associated with the site. It also states that no official closure has been announced, the original viral photograph cannot be reliably dated, and the duration of the disruption remains unknown.
That assessment may still be wrong. It is nevertheless more useful because its evidence, uncertainty and possible failure modes are visible. The decision-maker can act without being given a false impression of certainty.
The machine you no longer need
The most valuable feature of a reconstruction-first system may also be the hardest one to demonstrate on a product page.
You can show the event model, the evidence graph, the timeline and the changing confidence scores. You can compare an early assessment with what later became known. What is much harder to show is the infrastructure that no longer exists.
There is no visible achievement in choosing not to retain ten thousand copies of the same claim. No one sees the crawler that did not need to be repaired, the brittle integration that was never written or the vendor contract that did not have to be renegotiated.
The storage layer was not optimised; much of the redundant material was never stored. The collection architecture was not made infinitely scalable because part of the scaling problem was rejected rather than solved.
This matters because the amount of content that can be produced is now effectively unbounded. Any architecture whose costs rise directly with content volume has accepted an endless race. It must grow as quickly as people and machines can publish.
The number of meaningful events does not increase at the same rate. One event may generate a million posts, but it remains one event.
Before adding a new source, ask what unique evidence it contributes. Before signing another data contract, ask whether it provides a genuinely different view of the world. Before storing another billion posts, ask whether they represent a billion new facts or one fact repeated a billion times.
The intelligence industry has spent years becoming better at seeing what is being said. The next step is learning how to reconstruct what is actually happening, show the evidence behind that reconstruction and identify what would change the assessment.
Some of the best technical decisions are defined not by what a team was capable of building, but by what it understood well enough not to build.