GoldFynch

GoldFynch is a cloud-hosted e-discovery platform. You load in a corpus, cull it in steps down to what the assignment covers, and produce that selection. This page is about the second part: the culling, and the settings that decide whether your selection holds up when somebody checks it afterwards.

Why you capture everything and filter only afterwards, rather than selecting on site, is a separate question. That one is on Why an e-discovery platform. This page starts once that decision has been taken and the data is in.

First things first: the account is the weakest point

Set up the account with two-factor authentication before you go on site, not after. You are uploading somebody else’s confidential data to a cloud service, and from that moment access to that account is the narrowest point in the whole chain. Narrower than the service’s encryption, narrower than your own disk. Arrange it when you have the time to do it calmly, not in a corridor next to a custodian.

Loading in

A PST is compressed and expands considerably on ingestion. Reckon on roughly double. That is not a detail: with a platform that bills by case size, that figure decides which price band you land in. So estimate the volume after expansion and check the current rates before you create a case. How case creation and the upload itself work is covered in GoldFynch’s own documentation: creating a case and uploading files.

Some of the files cannot be processed. Empty files, or formats the platform cannot open. They come out of the ingestion as errors. Do not keep quiet about them: note how many there were, what you did with them, and include them in the dataset for completeness rather than letting them disappear quietly. A reader who cannot account for the difference between two counts will start looking for something behind it.

Cull in cumulative steps

Filter in steps, cumulatively, and give every step its own tag. For example:

  1. Folder scope. Inbox, sent items and deleted items. Deleted items do belong in there; what somebody threw away is often exactly what the assignment asks about.
  2. Date range. The period the court imposed.
  3. Keywords. The search terms from the assignment.

Report after every step how many files remain. That funnel with its counts is what allows a reader to check your selection. Without those counts there is a gap between “the complete mailbox” and “the documents produced”, and it is a gap you get asked about in a hearing.

The order of magnitude it comes down to in this kind of case: from tens of thousands of files to a few thousand. What the funnel mainly has to show is the rows between those two extremes, each with the step that accounts for the difference.

GoldFynch culling funnel Cumulative culling: folder scope, date range, keywords and finally families. From tens of thousands of files down to a few thousand. Without the "apply tag to the whole family" setting the attachments are missing. Folder scope tens of thousands of files Date range Keywords + families, tagged a few thousand files + families, untagged attachments missing: compare with the bar above Each bar cumulative on the one before. Width shows order of magnitude, not exact scale.

Attachments: one setting that makes the difference

Switch on apply tag to the whole family. Leave it off and you tag the message without what hung from it, and you end up with a selection whose attachments are missing, while those attachments are usually the very documents at issue. See tagging file families in the GoldFynch documentation for how that setting relates to the other two options (always ask, item only).

The difference in the counts is substantial and it has to be visible in your funnel. Put the step with and without families side by side, so a reader can see that you thought of it.

Keywords: this is where most goes wrong

The syntax with AND, OR, NOT and the remaining operators is set out in searching for keywords and advanced search in the GoldFynch documentation. What follows are the pitfalls the syntax itself does not mention.

  • Search on whole words. A keyword that also occurs in the custodian’s email domain otherwise matches a hundred percent of the corpus. The filter step then does nothing, and the only way you notice is by reading the counts.
  • Use full names, not surnames. Two people with the same surname on opposite sides of a dispute is nothing rare.
  • Search on the far side of the arrow. Searching a company’s mailbox for the name of that same company returns everything. Search on the counterparty, on the third party, on the file reference: on whatever sits at the other end of the relationship.
  • Test short terms before you commit to them. Anything under four or five characters is guaranteed to bring noise. Run such a term once with and once without surrounding spaces and compare the counts before you take it into your filter series.
  • Count on keywords matching file and folder names too, not just content. Sometimes that helps, sometimes it contaminates. Knowing which of the two it is means looking at the hits.
  • Define relative expressions of time explicitly. More recent than is read one way by one party and another way by the other. Write out in your report which date range you actually applied, in figures.

Produce in native format

Choose the native format, not PDF. Conversion destroys metadata that can turn out to be decisive later: send times, headers and embedded properties. At the moment of production you do not miss them, and that is exactly why it goes wrong. Whoever needs a readable version makes one afterwards from the original; the road back does not exist. The production settings, among them the choice between “Natives only”, “PDF only” and the load file or database format, are covered in productions in GoldFynch.

Redaction, and the trap behind it

Where the assignment has to protect a trade secret, there are two roads: exclude files entirely, or black out passages. When blacking out, choose the definitive mode (Final), not the preview mode (Preview). In preview mode the text is still underneath and it travels into the production. See redacting in GoldFynch for the distinction between the two modes.

The real trap sits deeper. A redacted child renders its container unusable. If the redacted file is inside a ZIP, that ZIP can no longer be produced in native format and you get a slip sheet in its place. The rest of the contents of that archive then drops out of your production. In that case deliver the extracted folder structure loose alongside it, and put in your report why.

The verification round

Load the finished production back in as a new, empty case and search it for precisely what should have been taken out. Zero results is the proof you hold in your hands before delivery instead of after.

This step costs a quarter of an hour and replaces heaps of good intentions. It is also the last moment at which a redaction error can still be repaired without consequences.

Closing down

Delete the data irrevocably from the platform as soon as the production has been accepted, and note the date and the method in your report. Noting it is the part that gets forgotten. A deletion recorded nowhere cannot be checked by a party, and as far as that party is concerned it might as well not have happened.

Edge formats

An Access database does not go in directly. Convert to CSV first and report how many tables that produced and how many records ultimately met the search criterion. That is the same funnel logic as with mail: the reader has to be able to work through from the source database to the rows produced.