SearchPractice

Keyword search misses coded language, which is the daily reality of narcotics work

A keyword list is a record of what one investigator knew on one day, and it starts decaying immediately.

Published 18 August 2026 6 min read Alexandre Ansart

Every narcotics unit has a keyword list. It was written by somebody good, it works, and it is out of date.

Not because anyone neglected it. Because the vocabulary it describes is actively adversarial. The people using it change terms when they suspect a term is known, they use different terms in different groups and different regions, and a large part of the meaning is carried by context rather than by any word at all.

This is the daily working reality of narcotics, trafficking and organised crime investigation, and it exposes something structural about how forensic search has been built.

What exact matching can and cannot do

Keyword search answers one question extremely well: does this exact string appear. That is genuinely valuable. It is fast, it is exhaustive within its scope, it is deterministic, and when you are looking for a phone number, an IBAN, a specific name or a known handle, nothing beats it.

The limitation is that it can only find what you already thought of. Every miss is silent. A search returning zero results looks identical whether the term is absent or whether it was present in a form you did not anticipate.

In practice, misses come from four directions. Synonyms and slang, which vary regionally and evolve. Deliberate substitution, where an ordinary word carries a specific meaning inside a group. Spelling, including transliteration, deliberate misspelling and autocorrect. And meaning carried entirely by context, where nothing in the sentence would ever appear on any list.

That last category is the one no keyword list can ever cover, and it is common. A message reading “same as last time, come alone” is potentially decisive and contains no searchable term at all.

What meaning-based search adds, and what it costs

Semantic search matches on meaning rather than on characters. A query for arranging a meeting retrieves messages about seeing each other, coming by, the usual place, or being around on Saturday, whether or not any word in the query appears in them.

The cost is precision. Meaning-based retrieval will confidently return things that are topically similar and investigatively irrelevant, and it has no notion of exactness, which makes it poor at precisely the queries keyword search is best at. Ask it for a specific phone number and you may get several plausible numbers.

So the two approaches fail in opposite directions. One misses things that are there. The other returns things that are not relevant. Choosing between them is a bad trade, and the reason it has been presented as a choice is mostly that the two were built as separate products.

Running both is the obvious answer

Running exact and semantic retrieval together and merging the results gives you exhaustiveness where the term is known and reach where it is not. It also gives the examiner something more useful than a single ranked list: a way to see which results came from where.

Two additions matter more than the fusion itself.

Query expansion. Automatically broadening the search to related vocabulary, so a query for one term also retrieves its variants, its slang forms and the substitutions used in its place. This is where an investigator’s accumulated knowledge belongs. Not in a static list that decays, but in expansion rules that can be maintained, versioned and shared between units.

Scope. The same query means different things at different levels. Across a whole case it is triage. Within one seal it is examination. Within a single conversation it is reading. An examiner needs to move between those without rebuilding the query.

Two things worth being careful about

Semantic search introduces a failure mode that keyword search does not have, and it deserves naming.

An exhaustive keyword search over a defined corpus can be described precisely in a report: this term, this scope, this many hits. A semantic search cannot be characterised the same way. There is no clean statement of what it did not return, and an examiner who says they searched for evidence of a meeting has made a much vaguer claim than one who says they searched for a specific string.

The practical answer is that semantic search is a tool for finding leads, not for making exhaustiveness claims. What it surfaces gets read, verified against the source and reported on its own terms. The path that found it is not the evidence.

The second is that a message meaning something in context is an interpretation, and interpretation is the examiner’s job and the court’s, not the software’s. A system that surfaces a thread because it resembles arranging a meeting has done something useful. A system that reports that a meeting was arranged has overstepped, unless it hands over the exact messages and lets a human decide.

That is why every result and every assistant statement should carry the record behind it, and why one click should put the examiner on the original row. The value of semantic reach is that it finds the thread nobody knew to look for. The verification is what makes that finding usable afterwards.

VERA

VERA is forensic analysis software for seized devices. It structures a raw Full File System extraction, makes it searchable and questionable with every answer cited, and runs entirely offline on your own infrastructure.

Request a demonstration

See it running on a real extraction

Thirty minutes, on fictitious data, with time for your questions. If VERA does not fit what your service needs, that is a useful answer too.

Every request is reviewed before access is granted.