Argus

Concepts

Review queue

Detections below the auto-confirm threshold go into a review queue at /review. Each card shows the detected face crop, the current match and similarity score, ranked suggestions from enrolled faces, and a View in image link that opens the full tag page with the detection bbox highlighted.

Three tabs

Key thresholds

All configurable in Settings.

Setting Default What it controls
face.match_threshold 0.5 Minimum similarity to assign a match at all. Below this, the face is stored but left unidentified.
face.auto_confirm_threshold 0.80 Detections at or above this are confirmed automatically and skip the queue. Below it, they land in the queue.
face.auto_enroll_threshold 0.92 Gate for the automatic enrollment path only. When Argus auto-confirms a high-similarity match, the embedding is added to the reference set only if the face-detection quality score clears this bar. Set to 0 to disable automatic enrollment.

Human actions always enroll

When you confirm, reassign, or label a face yourself, that's ground truth — its embedding is added to the person's reference set unconditionally. The auto-enroll threshold above does not apply to human actions.

Matching strategy

Controlled by face.match_strategy in Settings.

Keyboard shortcuts

Key Action
↑ / ↓ or W / SNavigate between cards
SpaceOpen / close the source-image zoom for the focused card
CConfirm the focused card
DDismiss (Suggested) / Unassign (No match, Mismatches)
VConfirm all in the focused card's group (Suggested only)
FReject all in the focused card's group (Suggested only)
AToggle select all on the active tab
Shift+CConfirm all selected
Shift+DDismiss / Unassign all selected

All shortcuts are suppressed when focus is in a text input.

Suggested people (face clustering)

The Suggested page (/clusters) groups unlabeled faces — ones that match nobody enrolled — into "probably the same person" clusters by similarity. Name a group and every face in it is labelled and enrolled together. This is the fast way to seed recognition on an existing photo set.

It complements the review queue: the queue handles faces that do resemble an enrolled person; clustering handles the residual unknowns that match no one yet.

Available over the API: GET /api/clusters?threshold=<0-1>&min_size=<n> returns the groups with detection ids and crop URLs. Read-only — clustering is computed on demand and stores nothing.

Manual face tagging

When the face detector misses someone — profile shots, occluded faces, poor lighting — you can draw a bounding box yourself on the tag page. Click-drag on desktop; long-press-drag on mobile. Label the box with a name and save.

Three-tier embedding fallback

After saving a manual box, Argus runs three attempts in order to extract a face embedding:

  1. Aligned (tier 1) — RetinaFace re-runs on the cropped region to detect and align the face, then ArcFace extracts an embedding from the aligned face. Best quality. Most manual boxes on a reasonably visible face land here.
  2. Unaligned (tier 2) — RetinaFace found nothing, so ArcFace runs directly on the raw crop without alignment. The embedding participates in matching but accuracy is lower. Useful for side profiles or partially covered faces.
  3. No embedding (tier 3) — ArcFace also returned nothing. The detection is saved as a labelled crop and appears in the identity's gallery, but it won't match against future detections. Rare — typically only very small or heavily obscured faces.

Border colors on the tag page

Manually drawn boxes use a dashed border. Color indicates tier:

ColorMeaning
GreenAligned embedding extracted (tier 1)
AmberUnaligned embedding extracted (tier 2)
RedNo embedding found (tier 3)

Auto-detected boxes use a solid white border regardless of outcome.

Object detection models

Standard YOLO

Models: yolov8n, yolov8s, yolov8m, yolov8x, yolo11n

Detects a fixed set of 80 everyday object categories defined by the COCO dataset (people, vehicles, animals, furniture, food, etc.). Fast and consistent — the vocabulary is baked into the model weights. Use the Object classes setting to filter which of the 80 you care about.

YOLO-World (open vocabulary)

Models: yolov8s-worldv2, yolov8m-worldv2, yolov8l-worldv2

Detects anything you describe in plain language. Instead of a fixed list, you define a vocabulary of words and phrases, and the model finds those things in photos.

YOLO-World understands natural language, so descriptions like "golden retriever", "broken window", or "person on a bicycle" work. Abstract or non-visual concepts do not — if you couldn't point at it in a photo, it won't work.

Argus ships with ~160 default classes covering all 80 COCO categories plus common additions: weapons, fire, smoke, license plates, face masks, extended vehicle types, more animal species, and other frequently useful categories. Edit the vocabulary in Settings → YOLO-World vocabulary. Changes take effect on the next detection — no restart needed.

Use standard YOLO when the things you want to detect are within COCO's 80 classes. Switch to YOLO-World when you need to detect things outside that list.

Environments

Each user has one or more named environments. All recognition data — identities, detections, enrolled faces, source images — is isolated per environment. Switching environments is instant; data in other environments is never visible.

Use cases:

Manage environments at Account → Manage environments. Deleting an environment permanently removes all its data.

API keys are environment-scoped. Each key is bound to one environment at creation time. Requests authenticated by that key read and write only that environment's data, regardless of which environment the browser session has active.