Adapted from the OsintLLM repository docs. Private repo. PROPRIETARY — do not make public.
An experimental local-first OSINT research platform. The goal: a specialized local model and software framework that operates as an intelligent command center for open-source intelligence work — planning investigations, calling trusted open-source tools, preserving evidence, connecting entities, generating reports, and assisting analysts with detailed public-source research.
Status: concept and research phase. Treat everything here as R&D direction, not shipped product.
A self-directing OSINT command center built around a forked or fine-tuned open-source model:
- Understand a research objective, break it into safe/lawful/auditable steps
- Select and run the right open-source tools per step (locally or via controlled adapters)
- Preserve source material and evidence trails; extract entities (people, orgs, usernames, domains, locations, vehicles, documents, events…)
- Build timelines, maps, link graphs, evidence packets; compare claims against archived material
- Generate analyst-ready cited reports; suggest/build custom tools when existing ones fall short
- Open-source first — open tools, data standards, forkable infra; proprietary APIs only as optional connectors. Tool inventory:
docs/open-source-osint-tools.md.
- Local-first — local model execution, local evidence stores, analyst-controlled config; clear boundaries on what leaves the machine.
- Analyst-in-the-loop — assists, doesn't replace judgment; every finding traceable to source material, tool output, timestamps, reasoning notes.
- Evidence preservation — original URLs, archive URLs, captures, files, metadata, hashes, analyst notes, reliability assessments. Always distinguish source claim vs model inference vs analyst conclusion vs verified fact vs unresolved lead.
- Modular tool orchestration — tools wrapped in adapters with clear contracts (name, license, I/O schemas, safety restrictions, rate limits, evidence produced). No arbitrary shell calls.
- Safety/legality/auditability — lawful ethical public-source research only. Explicitly not: doxxing, harassment, stalking, credential theft, phishing, malware, unauthorized access, private-data acquisition.
Analyst UI (cases, tasks, evidence, timelines, maps, graphs, reports)
│
OsintLLM orchestration core (planning, tool routing, safety, memory, citations)
│
Local model layer (open-source fork/fine-tune/quantized runtime)
│
Tool adapter layer (Bellingcat, crawlers, geo, entity, archive, APIs)
│
Data layer (evidence store, entity graph, vector DB, search, SQL/DuckDB)
│
Audit & reporting layer (logs, source trails, hashes, exports)
Suggested stack: Rust/Python/TypeScript core · llama.cpp/Ollama/vLLM · Python-first adapters · PostgreSQL/SQLite + DuckDB · OpenSearch/SQLite FTS + Qdrant · FollowTheMoney-inspired entity schema.
- Phase 0 — research & architecture: tool inventory ✅, scope/ethics policy, model/runtime choice, adapter/evidence/graph/case schemas
- Phase 1 — local prototype: app shell, cases, evidence storage, URL archiving, ingestion, search, planning loop, audit log
- Phase 2 — tool orchestration: Bellingcat/Wayback/username/domain/document/geo adapters, output normalization
- Phase 3 — entity graph & reports: extraction, dedup, timelines, cited report drafting, exports
- Phase 4 — advanced: fine-tuned behavior, cross-case matching, geolocation/media-verification workflows, custom tool builder
Strategic note: the roadmap recommends reusing Lumen aggressively (agent runtime, approvals, audit, plugin system) rather than rebuilding it.
Investigative journalism, public-interest research, corporate due diligence, sanctions/entity research, threat intel, misinformation analysis, conflict monitoring, geolocation/chronolocation, public records, media verification, web archives, infrastructure mapping (authorized scope), evidence collection and reporting.
- Repo:
black-candle-technologies/OsintLLM (private, proprietary)
- Related: Lumen (recommended runtime reuse)