Why an e-discovery platform, and not filtering on site
Sooner or later the opposing party asks the question: why did you take the whole mailbox, when the assignment covered a limited period and a handful of keywords? The question is fair and it deserves a technical answer. Anyone who falls back on “that is simply how it works” has already lost the argument.
The answer is that selecting on site cannot be done in a way that holds up. There are three reasons for that. They are independent of one another and each of the three is sufficient on its own.
An email client does not export a partial selection
The built-in search of an email client lets you search and inspect the hits, but it cannot write those hits out as a separate set. Export runs per mailbox: all or nothing. There is nothing in between, where the filtered set goes out and the rest stays put.
Anyone determined to settle it on site has to move the items found into a separate folder by hand. That is a write operation in the custodian’s mailbox: after your visit the system is no longer what it was when you walked in.
Re-indexing takes longer than your visit
Searching presupposes an index, and at the moment you need one it is usually not there. An exported or copied mailbox opened in an email client starts from zero. Rebuilding that index easily takes twelve to twenty-four hours, depending on the volume and on the disk underneath. That is more than a full working day on site.
The trap behind it is worse than the waiting. A search against an index that is still building returns incomplete results, and the search function does not report that it is unfinished. The result looks like a success. Whoever decides at that moment what comes along and what stays behind takes that decision on half a dataset without knowing it.
Content search depends on the format
There is no single search that looks inside PDF, Word, PowerPoint and plain text at the same time. Every format needs its own text extraction, and where that extraction is missing the search silently falls back to file names and metadata. A corpus made up of all those formats mixed together, which is the normal situation on a company share, cannot be searched in full on content with the means present on the machine itself.
The alternative is opening every file in the application that can handle it. With a corpus of tens of thousands of files that is out of reach in a single visit. This is not a matter of working harder.
Why filtering afterwards is more accurate
Whoever selects on site takes decisions that are recorded nowhere: standing next to someone else’s desk, under time pressure, on the basis of what the search happened to show at that moment. Afterwards there is no way to establish why one file was in the selection and another was not, and the files you left behind cannot be assessed a second time.
In an e-discovery platform every filter step is a separate operation with a count before it and a count after it. Folder scope, date range and keywords are applied one after the other and every intermediate state survives. That funnel is traceable: an opposing expert who takes the same steps arrives at the same counts. Collecting everything and then filtering in a way you can account for is therefore not sloppier than selecting on site. It is more accurate, because the judgement becomes reproducible instead of final.
What over-collection obliges you to do
The choice is defensible only with the closing piece attached. Without that closing piece you have merely taken more than was asked.
What falls outside the assignment is not produced. The fact that you had it in your hands gives you no ground whatsoever to deliver it. Borderline cases you do not settle yourself: you put them to the court that gave the assignment.
The source data is irreversibly wiped when the work is done. From the platform and from your own media. Record when, where and by what method, and put that in the report. As long as the destruction has not been shown, the collection stays an open risk for the custodian, and one he can do nothing about himself.
Notify all parties, including when there is nothing to see
Notify all parties of a data collection in writing, including when the step looks purely technical and there is nothing of substance to discuss. The right to tegenspraak, the adversarial principle that runs through the whole Belgian expert investigation and requires every party to be able to follow and contest each step, does not depend on whether there is anything to discuss. It attaches to the operation itself.
A collection is moreover the one moment in the investigation that cannot be done over. The state of the system at that instant does not come back. Anyone who has to explain afterwards why a party was not there is no longer arguing about his method but about his impartiality.
Further reading
Craig Ball, Don’t BE a Tool, GET a Tool! (22 March 2021). Ball comes at the same question from the integrity argument: assessing evidence in the application that normally opens the file changes metadata and breaks the hash value. His conclusion runs parallel to the one above, by a different route.