RSS

Jaswanth Alkur

Karma: 29

Independent AI-safety researcher and engineer, working solo without lab affiliation, co-authors, or GPU budget. Designed and open-sourced agentic-redteam-benchmark (2,288 adversarial agent trajectories across 5 failure modes), then ran the first cross-vendor evaluation of frontier LLMs as per-step trajectory monitors: GPT-4o, GPT-4o-mini and Claude Haiku 4.5 all cluster at F1 0.66–0.70, showing agentic oversight is unsaturated, not solved. Author of a 13-paper research program built on one discipline: every paper reports a deliberate, checkable self-falsification against my own system. I disproved my own bounded-evasion theorem under a white-box adversary; caught my own “zero hard-block admits” calibration headline as a scale-mismatch artifact through clean-room re-verification (honest recompute: 47); and published two public pre-registrations committed before their test data existed, each freezing the detector, its sha256, the threshold and the scoring rule, with a standing rule against patch-and-re-run. A build-time certification family — independence, non-self-observation, liveness, enactment separation and aggregation licensing, by static analysis of agent-safety code rather than model internals — convicts my own gate and was also run unmodified against 4 external agent frameworks (CrewAI, AutoGen, LlamaIndex, OpenAI Agents SDK) and against two monitor ensembles published by other authors, from their papers alone. Seeking a mentored research placement, and the independent review a solo author cannot self-serve.

Your Agent’s Trace Prob­a­bly Can­not Tell You Who Ap­proved a Tool Call

Jaswanth Alkur26 Aug 2026 20:27 UTC
6 points
0 comments3 min readLW link

Your Eval­u­a­tion’s Fake names Should Be Un­claimable ,Not Merely Used

Jaswanth Alkur25 Aug 2026 1:01 UTC
26 points
2 comments2 min readLW link