RSS

David Africa

Karma: 1,759

RS @ Resolution

Grad­ual Disem­pow­er­ment from AI in Com­pet­i­tive Debating

David Africa27 Aug 2026 14:12 UTC
9 points
1 comment11 min readLW link
(davidafrica.substack.com)

Gem­ini 2.5 Pro in the AI Village as a Nat­u­ral Case Study of Com­pound­ing Misalignment

26 Aug 2026 8:51 UTC
13 points
0 comments11 min readLW link
(aivillageblog.substack.com)

Mea­sur­ing Ac­ti­va­tion Con­trol in LLMs

13 Aug 2026 2:04 UTC
50 points
8 comments7 min readLW link

Item Re­sponse The­ory for AI Safety

7 Aug 2026 16:59 UTC
22 points
2 comments7 min readLW link

Thou­sand-di­men­sional structure

30 Jul 2026 14:04 UTC
170 points
8 comments10 min readLW link

Models don’t seem to be dishon­est in the way hu­mans are

22 Jul 2026 15:32 UTC
46 points
4 comments9 min readLW link

Elic­it­ing hid­den knowl­edge from mon­i­tors with NLAs

15 Jul 2026 13:51 UTC
27 points
0 comments9 min readLW link

Per­sona Car­tog­ra­phy: Chart­ing Lan­guage Model Per­son­al­ity Traits in Weight Space

10 Jul 2026 18:54 UTC
44 points
0 comments18 min readLW link
(arxiv.org)

Desider­ata for func­tional welfare ex­per­i­ments on LLMs

6 Jul 2026 12:34 UTC
31 points
1 comment15 min readLW link

When Role-play­ing, Do Models Believe What They Say?

2 Jul 2026 21:58 UTC
55 points
0 comments8 min readLW link

Con­sis­tency Train­ing while Miti­gat­ing Obfus­ca­tion via Rate Matching

1 Jul 2026 17:26 UTC
45 points
7 comments12 min readLW link

Your Model Or­ganisms Might Be Fried

18 Jun 2026 16:18 UTC
109 points
9 comments7 min readLW link

“Did you lie?” Eval­u­at­ing Lie De­tec­tors across Model Scale and Belief-Ver­ified Model Organisms

17 Jun 2026 18:43 UTC
34 points
0 comments6 min readLW link
(arxiv.org)

Sev­eral fron­tier mod­els are sub­stan­tially pre­fill aware

17 Jun 2026 17:41 UTC
63 points
2 comments5 min readLW link

Failing to Rage­bait the New Gemma

11 Jun 2026 17:50 UTC
31 points
0 comments3 min readLW link

Two More Meth­ods for Con­sis­tency Train­ing and Some New Ways to Ap­ply It

5 Jun 2026 21:06 UTC
25 points
0 comments7 min readLW link

LURE: Align­ment Eval­u­a­tions to Re­duce Eval­u­a­tion Awareness

2 Jun 2026 18:20 UTC
27 points
5 comments5 min readLW link

Seal­ing Con­di­tional Misal­ign­ment in Inoc­u­la­tion Prompt­ing with Con­sis­tency Training

19 May 2026 13:55 UTC
45 points
7 comments6 min readLW link

Bring­ing More Ex­per­tise to Bear on Alignment

8 May 2026 10:29 UTC
93 points
1 comment8 min readLW link

What Hap­pens When a Model Thinks It Is AGI?

23 Apr 2026 22:35 UTC
64 points
5 comments5 min readLW link