Independent AI-safety researcher and engineer, working solo without lab affiliation, co-authors, or GPU budget. Designed and open-sourced agentic-redteam-benchmark (2,288 adversarial agent trajectories across 5 failure modes), then ran the first cross-vendor evaluation of frontier LLMs as per-step trajectory monitors: GPT-4o, GPT-4o-mini and Claude Haiku 4.5 all cluster at F1 0.66–0.70, showing agentic oversight is unsaturated, not solved. Author of a 13-paper research program built on one discipline: every paper reports a deliberate, checkable self-falsification against my own system. I disproved my own bounded-evasion theorem under a white-box adversary; caught my own “zero hard-block admits” calibration headline as a scale-mismatch artifact through clean-room re-verification (honest recompute: 47); and published two public pre-registrations committed before their test data existed, each freezing the detector, its sha256, the threshold and the scoring rule, with a standing rule against patch-and-re-run. A build-time certification family — independence, non-self-observation, liveness, enactment separation and aggregation licensing, by static analysis of agent-safety code rather than model internals — convicts my own gate and was also run unmodified against 4 external agent frameworks (CrewAI, AutoGen, LlamaIndex, OpenAI Agents SDK) and against two monitor ensembles published by other authors, from their papers alone. Seeking a mentored research placement, and the independent review a solo author cannot self-serve.
Jaswanth Alkur
Karma: 29