Digercules

Digercules is a program with neural-network components that takes any raw, unsorted pile of files — an archive, a library, concert recordings, messenger correspondence — and turns it into a structured library, including smart deduplication: finding every copy of the same thing, removing the excess, and losing nothing that matters. For every single file — a topic, a language, a document type, key entities, a short summary. This is done ahead of time, without rush or deadline, before anyone urgently needs it — not at the last minute under pressure for a specific task. The finished library can be fed to any neural network, database, or search system — they take it from there, finding what's needed and building the interface; Digercules doesn't replace that next step, it makes it possible in the first place, by taking on the labor-intensive preparation nobody ever has the resources for.

Data is processed without corporate clouds — either entirely on the client's own computer (given adequate GPU capacity), or the program's author is sent, not the source files, but compressed copies prepared specifically for recognition, with only text coming back.

Already proven on archives of varying scale — real statistics from processed projects:

800,000+
files processed
~85 TB
source volume
66,000+ hrs
of audio (~7.5 years)
1.8 TB
reclaimed via deduplication
3,507 hrs
spent processing
~430
performers identified on video

Over the available connection (~30 Mbit/s, measured) uploading this much data to the cloud would take about 262 days — versus the 96 it actually took to process on-site. Storage would run $84/month, and processing on a comparable GPU would run $150–190/month. And that's without the next step: the data still has to move from storage onto the GPU machine itself for processing — even over a connection three times faster (100 Mbit/s), a batch of a few terabytes takes about 3 days.

What's Actually Running Under the Hood

Digercules isn't one model — it's an orchestrated pipeline of 25+ processing phases, each doing one narrow job well, chained so the output of one becomes the input of the next.

Speech to text. Audio and video go through faster-whisper (large-v3-turbo): a fast first pass for matching and triage, then a slower, careful main pass with guided prompts built from the archive's own vocabulary. When an archive has its own recurring cast of voices or recurring terminology, the model can be fine-tuned (LoRA) on that specific corpus — proven on a hundred-plus hours of a single collection.

Voice identification. A speaker-embedding model (ECAPA-TDNN) builds a voice gallery — who's speaking, matched against everyone the system has already heard, growing with every archive it processes.

Face identification. The same pattern as voice identification, applied to faces: a multi-prototype gallery built to survive what a single averaged face can't — recognizing the same person photographed as a child and forty years later. The confirmation screen has been proven on 20 TB of real concert video footage.

Face review screen: confirmed cluster gallery for one person
Review screen: every confirmed appearance of one person's face, grouped by video.
Face review screen: suggested matches with similarity percentage
The same screen mid-review: the system suggests a match with a similarity score, a person confirms or rejects it.

The decision is always the person's: similarity is a hint with a percentage attached, not a finished answer, and there's no safe threshold past which it can be trusted automatically — the interface doesn't hide that. A face nobody recognizes isn't passed off as identified; it waits for whoever knows the archive better. More on how the module works.

OCR. Scanned documents go through Marker — an open-source PDF/scan recognition engine — turning a photographed page into searchable text, not just a picture of one.

Photo understanding, in tiers. Every photo bank first gets a cheap, CPU-only pass (OpenCV DNN — EAST for text regions, YuNet for face counts) that runs identically on every machine in the fleet, including decade-old Windows 7 hardware. On newer hardware, a second pass adds CLIP zero-shot scene classification at GPU speed. Where it matters most, a full vision-language model (Ollama) writes an actual per-photo description. Three tiers, cheapest first.

Document understanding. Every file gets read and described by a local LLM (Ollama): what it's about, in what language, what it's really for.

Deduplication that also protects your files. The same hash that finds duplicate copies also verifies, byte for byte, that a file arriving back in your archive after processing is the exact file that left it.

Dense wall of hundreds of recognized face variants from one archive
Scale, on one real archive: hundreds of recognized face variants, auto-grouped and human-refined.

Every one of these runs locally or on a specific, named machine you choose — never a faceless cloud API.

Status

The product is in active development. The engineering library is stable, in production.

Team Geography

Development is done by a distributed team across several countries: Israel, Georgia, Russia. Real points of presence for the compute infrastructure: Rishon LeZion, Israel · Tbilisi, Georgia · Chelyabinsk, Russia · Moscow, Russia.

Where to Next

Digercules serves different types of customers through separate sites:

How This Is Priced

A local, self-run pass is a one-time purchase of the tool with no cap on archive size; only the preliminary estimate, from a description or screenshot of the archive, is free. Heavier coordination, when your own hardware isn't enough, is separate and priced by the actual volume of work found — there's no fixed number for that yet. For a market reference, see the figures earlier on this page (what a comparable job would cost through an ordinary cloud). Details for each scenario are on digercules.org and digercules.net.