MediaPractice

Media volume in child protection cases: what triage can and cannot do

This subject deserves to be written about carefully or not at all, so this piece is deliberately narrow: what the arithmetic is, and what automation does and does not change.

Published 25 August 2026 6 min read Alexandre Ansart

This is a difficult subject to write about from a vendor position, and there is a version of it that should not be written at all: the version that treats the suffering in these cases as a marketing opportunity, or that claims detection accuracy nobody can evidence.

So this piece is deliberately narrow. It is about the arithmetic of media review, what automated triage genuinely changes, and what it does not.

The arithmetic

A single seized device commonly holds tens of thousands of images and video files. A case with several devices reaches six figures without being unusual. Investigations involving distribution networks can involve substantially more.

Human review of that volume is slow by necessity. It is slow because it requires attention, because it requires judgement about material that is often ambiguous at the edges, and because the people doing it cannot sustain it for long periods. Well-run units limit exposure time deliberately. That limit is not inefficiency, it is an occupational health control, and any calculation that treats reviewer hours as elastic is wrong before it starts.

So the constraint is not the volume by itself. It is volume against a supply of review hours that is deliberately and correctly capped.

What automation actually changes

Content-based indexing of media allows the collection to be searched by what it contains rather than only by filename, timestamp or hash. That changes three things concretely.

Ordering. A reviewer working through material in an order informed by content is more likely to find what matters early. In cases where identifying a victim quickly changes what happens next, ordering is not an efficiency gain, it is an outcome.

Elimination. Large parts of a media collection on any device are wallpapers, application caches, memes, screenshots and downloaded assets. Removing them from the review queue is low risk and it is a substantial fraction of the volume.

Linking. Attachments tied back to the conversations that carried them, and deduplication against material already seen, turn a flat pile of files into something with structure. Knowing which thread an image arrived in is frequently more informative than the image.

Two of those three do not require the software to make any judgement about the content of the material at all, which is worth noting, because they carry most of the practical benefit.

What automation does not change

It does not remove the examiner from the loop, and it does not reduce the standard of what a court requires.

Any system that classifies material is operating in a domain where the categories are legally defined, jurisdiction-specific, and contested at the boundaries. Age estimation from an image is not reliable in the range that matters most. Context frequently determines the character of an image in ways nothing in the pixels indicates.

A vendor claiming an accuracy figure for this task is making a claim that cannot survive cross-examination, because the benchmark it came from is not the population of a real case, and the categories it used are not the ones the statute uses.

So the honest framing is that automated triage changes the order and the size of the review queue. The determination of what material is remains a human and a legal judgement, and the software’s job is to make the human’s work possible rather than to pre-empt it.

Hash matching, and its limits

Known-file hash matching remains the most reliable tool for this work and the least discussed, probably because it is unglamorous. Where material matches a known set, the identification is exact and defensible in a way no learned system is.

Its limitation is equally clear. It only finds what has been seen before and catalogued. New material, re-encoded material and anything produced locally will not match, and locally produced material is often the most urgent, because it may indicate ongoing harm.

That is the gap content-based retrieval addresses, and it is worth being precise about the claim. It does not identify material as illegal. It surfaces material that resembles a description, so that a human looks at it sooner than they otherwise would.

What to ask a vendor

If you are evaluating any tool that claims media triage capability, the useful questions are narrow.

Ask what happens to the material: whether any of it leaves your infrastructure at any point, including for indexing. For this category of case, an answer involving a third-party service is usually the end of the conversation, and it should be.

Ask what the system claims to determine, and what it merely ranks. Those are different, and vendors blur them.

Ask what the reviewer sees. A tool that presents a decision invites the reviewer to confirm it. A tool that presents ordered material and lets the reviewer decide is doing the same arithmetic without borrowing the reviewer’s authority.

Ask about exposure. Whether the interface can present material in a way that reduces unnecessary exposure while remaining sufficient for the judgement being made is a real product question, and one that units care about more than most vendors expect.

The commitment

VERA indexes media so it can be searched by content, and the assistant can be asked to look for specific elements. All of it runs on the customer’s own infrastructure, offline, and no material leaves the network at any point.

We do not publish a detection accuracy figure, and we will not, because we cannot evidence one in front of you on your own data. If that is what your evaluation requires, we would rather say so at the start than be found out in the middle.

VERA

VERA is forensic analysis software for seized devices. It structures a raw Full File System extraction, makes it searchable and questionable with every answer cited, and runs entirely offline on your own infrastructure.

Request a demonstration

See it running on a real extraction

Thirty minutes, on fictitious data, with time for your questions. If VERA does not fit what your service needs, that is a useful answer too.

Every request is reviewed before access is granted.