Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
David Africa
Karma:
1,405
Research Scientist with the Alignment team at UK AISI.
All
Posts
Comments
New
Top
Old
Page
1
Eliciting hidden knowledge from monitors with NLAs
David Africa
and
Aleksandr Bowkis
15 Jul 2026 13:51 UTC
27
points
0
comments
9
min read
LW
link
Persona Cartography: Charting Language Model Personality Traits in Weight Space
antonghawthorne
,
Mariia Koroliuk
,
Irakli Shalibashvili
,
sidbaines
,
Clément Dumas
,
Konstantinos Voudouris
and
David Africa
10 Jul 2026 18:54 UTC
42
points
0
comments
18
min read
LW
link
(arxiv.org)
Desiderata for functional welfare experiments on LLMs
Rikhil Jhaveri
,
Jamie Johnson
and
David Africa
6 Jul 2026 12:34 UTC
30
points
1
comment
15
min read
LW
link
When Role-playing, Do Models Believe What They Say?
Sturb
,
David Africa
and
Sid Black
2 Jul 2026 21:58 UTC
54
points
0
comments
8
min read
LW
link
Consistency Training while Mitigating Obfuscation via Rate Matching
Sohaib Imran
,
Prakhar Gupta
,
Jannes Elstner
and
David Africa
1 Jul 2026 17:26 UTC
45
points
7
comments
12
min read
LW
link
Your Model Organisms Might Be Fried
Daniel Tan
,
J Bostock
,
draganover
,
ma-rmartinez
,
sidbaines
and
David Africa
18 Jun 2026 16:18 UTC
102
points
9
comments
7
min read
LW
link
“Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Alan Cooney
,
David Africa
and
Geoffrey Irving
17 Jun 2026 18:43 UTC
34
points
0
comments
6
min read
LW
link
(arxiv.org)
Several frontier models are substantially prefill aware
yeedrag
,
Parv Mahajan
,
David Africa
,
alexsouly
,
Jordan Taylor
and
RobertKirk
17 Jun 2026 17:41 UTC
61
points
2
comments
5
min read
LW
link
Failing to Ragebait the New Gemma
Neil Shah
,
David Africa
and
arav-dhoot
11 Jun 2026 17:50 UTC
30
points
0
comments
3
min read
LW
link
Two More Methods for Consistency Training and Some New Ways to Apply It
David Africa
,
Sukrati_Gautam
,
Neil Shah
and
arav-dhoot
5 Jun 2026 21:06 UTC
25
points
0
comments
7
min read
LW
link
LURE: Alignment Evaluations to Reduce Evaluation Awareness
Igor Ivanov
and
David Africa
2 Jun 2026 18:20 UTC
27
points
5
comments
5
min read
LW
link
Sealing Conditional Misalignment in Inoculation Prompting with Consistency Training
David Africa
,
Sukrati_Gautam
and
Neil Shah
19 May 2026 13:55 UTC
44
points
7
comments
6
min read
LW
link
Bringing More Expertise to Bear on Alignment
Edmund Lau
,
Geoffrey Irving
,
Cameron Holmes
and
David Africa
8 May 2026 10:29 UTC
87
points
1
comment
8
min read
LW
link
What Happens When a Model Thinks It Is AGI?
josh :)
and
David Africa
23 Apr 2026 22:35 UTC
64
points
4
comments
5
min read
LW
link
Gemma Gets Help: Mitigating Frustration and Self-Deletion with Consistency Training
David Africa
and
Neil Shah
20 Apr 2026 16:07 UTC
27
points
1
comment
12
min read
LW
link
From personas to intentions: towards a science of motivations for AI models
David Africa
and
Jacob Pfau
14 Apr 2026 12:26 UTC
81
points
5
comments
7
min read
LW
link
Emergent stigmergic coordination in AI agents?
David Africa
15 Mar 2026 12:30 UTC
49
points
2
comments
3
min read
LW
link
Steering Awareness: Models Can Be Trained to Detect Activation Steering
josh :)
and
David Africa
12 Mar 2026 23:34 UTC
20
points
0
comments
6
min read
LW
link
Prefill awareness: can LLMs tell when “their” message history has been tampered with?
David Africa
,
alexsouly
,
Jordan Taylor
and
RobertKirk
9 Mar 2026 10:47 UTC
86
points
11
comments
10
min read
LW
link
A Proposal for TruesightBench
David Africa
5 Feb 2026 14:33 UTC
14
points
0
comments
4
min read
LW
link
Back to top
Next