Public demo: running on recorded data, no live API calls.
AGENT CENTRAL
LLM today: --
Jobs pipeline
Loading pipeline...
Welcome to Agent Central
A live control room for the real AI agents I build and run.
Every character is driven by real backend state over a WebSocket. When an
agent actually works, it walks to the matching room. Nothing here is faked.
Tip: click any agent, room screen, or object to inspect it.
The strip on the left and the ticker below update live.
Zones
Controllive ops & stats
Browsescanning the web
ThinkingLLM analysis
Libraryresearching
Writingdrafting
Commsposting & replying
Loungeidle / off-shift
Agents · live
LIVE
connecting to the live activity feed...
CONTROL · METRICS
Loading metrics...
SECRETARY
JOB SCOUT
JOB ANALYST
EVALUATOR · SCORECARD
This scorecard is built to survive a skeptical reviewer. Ground truth
comes from one of two places, and every check is labelled with which:
COMPUTED FROM ACTIVITY_LOG (objective)
The Secretary groundedness questions have answers computed directly
from the SQLite activity data with code. The database is the ground
truth, never a model's opinion. We also verify that every source the
Secretary cites is a real window with real events; a cited window that
does not exist is counted as a hallucination.
MODEL-DRAFTED (OPUS 4.8), CLEAR-CUT
The Job Analyst cases were drafted at full effort by a stronger model
(Opus 4.8) judging a weaker one (Haiku). The ground truth is only the
expected CATEGORY / band (high / low / filtered / deduped), never a
precise model-picked number. Cases are kept clear-cut; any that are not
obviously clear-cut are flagged needs review
for a human. A case a human has checked is relabelled
human-reviewed.
A same-tier model's subjective fine-grained score is
never treated as truth.