Turning fragmented information into source-grounded intelligence.
Most demonstrations of AI-assisted research show a model answering a question. The difficult part of doing this reliably is not the answering — it is everything around it: finding material, normalising it, keeping track of where each statement came from, and refusing to confuse repetition with confirmation.
This article is about that surrounding system. It is not about model capability; model behaviour is evaluated separately in Local AI Benchmarking on Real Infrastructure. Here we examine the engineering that makes a model’s capability usable.
Implemented as ARGUS, the Scelaris research pipeline.
1. Why Research Is Harder Than It Looks
Asking a capable model to "research a topic" produces text that looks like research. Whether it is research depends on properties that the model, left alone, does not provide.
Real information spaces are fragmented. The same underlying event is reported by several sources, in different formats, with different terminology and different levels of reliability. A naive system treats each report as an independent finding — so repetition quietly becomes corroboration. Formats drift over time; provenance dies at the point where material is copied out of its original context; and relevance is a moving target that depends on the operational question, not on the document.
None of these are model problems. They are information-space problems, and they are solved — or not — before the model sees anything at all.
2. Research as an Engineering System
ARGUS treats research as a pipeline of engineering disciplines, with language models as one component among several. Most of the work is deliberately deterministic.
| Layer | What runs | Model involvement | Preserved for evidence |
|---|---|---|---|
| Acquisition | Deterministic collection across source types; scheduling and retrieval. | None. | Original source, retrieval time, location. |
| Normalization | Deterministic extraction and canonicalisation: formats, units, terminology, encoding. | None. | Normalised form and the original. |
| Deduplication | Deterministic identity and similarity handling across reports of the same underlying event. | Advisory at most. | Grouping decisions, source set. |
| Correlation & evaluation | Context assembly, conflict presentation. | Semantic interpretation: what is present, what matters, what conflicts. | Which sources informed which statement. |
| Reporting | Structured generation with mandatory source references. | Composition under strict constraints. | Statement-level provenance. |
The deliberate design decision is visible in this table: the system keeps provenance bookkeeping, deduplication and normalization away from the model. A language model is excellent at interpreting language and unreliable at bookkeeping — so the system never asks it to keep the books.
This also defines where AI helps and where it does not. Acquisition does not need a model. Normalisation should not have one. Interpretation without normalized, deduplicated material is not intelligence — it is confident improvisation over unverified input.
3. Evidence as a First-Class Output
The output of a research system is only as trustworthy as its relationship to its sources. Two sentences can carry identical wording and completely different epistemic weight:
“The source states X.” — a fact about a document.
“The system concludes X.” — an interpretation that must be independently justifiable.
ARGUS keeps these explicitly separated. Every factual statement in a report carries a resolvable source reference; conclusions are distinguishable from citations. This traceability is not a reporting nicety — it is the property that makes the output reviewable, auditable and correctable by a human.
Deduplication deserves its own attention here. When four reports describe the same event, a naive pipeline produces four citations and the impression of four independent confirmations. In reality there is one event and one fact. Detecting that the same underlying occurrence has been reported repeatedly — and recording the whole source set behind it — is one of the quiet capabilities that separates a research system from a retrieval demo.
It is also where much of the engineering lives: identity decisions that are deterministic where possible, model-assisted where necessary, and always recorded.
4. ARGUS in Practice
ARGUS is currently operating on real research workloads. The following evidence is being captured from the running system and will be added here as it becomes available:
- Corpus shape: source documents, source types, volume and coverage.
- Processing funnel: acquired → normalized → deduplicated → relevant.
- Provenance coverage: share of report statements carrying resolvable source references.
- Processing distribution: deterministic vs. LLM-mediated steps.
- One end-to-end example: from source fragments to a cited report excerpt.
- Human review: what the reviewing engineer sees and corrects.
- Limitations and failures: sources that resisted normalisation, relevance errors caught in review.
- Maintenance behaviour: what changes between repeated runs over time.
Status: evidence pending. ARGUS is currently producing the real material intended for this section. Nothing is reported here until measurements have been reviewed — consistent with how Scelaris handles results: measured first, published after review.
Related Work
The model capabilities that ARGUS employs — and their measurable differences on real hardware — are evaluated separately in Local AI Benchmarking on Real Infrastructure. That article asks how well models perform research tasks. This article asks how a system around them makes that performance reliable.