April 23–27 · 5,471 papers reviewed · 28 flagged for you
Heaviest day — the mech interp and SAE cluster lands here. Two sessions back-to-back.
Your poster day. Session 3 in the morning, then you present in Session 4.
Post-poster day. Free to roam. Safety/alignment cluster lands here, plus Potts' causal interventions oral.
Rooms 201-204 (second floor). Each talk is ~12 min. Orals run parallel to poster sessions.
Nothing else in this block -- head to Poster Session 1 in Pav 4 for the SAE cluster (Temporal SAEs poster, AbsTopK, Amortized Inference, Paulo/Belrose) after this talk.
Recommended route: Start in 203 A/B for the 2D session (3:15-3:50ish). Dash to 204 C for Temporal SAEs at 4:03. If you want, return to 203 A/B for the 4:15 talk. Poster Session 2 runs in parallel in the pavilions -- hit FaithCoT-Bench and Thought Branches (Nanda) before or after the oral block.
Recommended route: Go to 201 A/B for GEPA at 11:30. When it ends (~11:42), walk to 202 A/B for the Liang/Hashimoto talk. Two Stanford networking opportunities back to back. Poster Session 3 runs in the pavilions (RFEval, GRPO Unlearning).
Only one relevant oral this block -- you'll mostly be at your poster (Pav 4, P4-#4306). Your poster neighbors (In-Context Algebra, Markovian Transformers, Nanda's planning paper, unlearning papers) are the priority.
Recommended route: Start in 202 A/B for RL Scaling at 11:06. At 11:18 (when it ends), move to 203 A/B for the Potts talk at 11:30. This is your best networking moment of the conference. Poster Session 5 has Circuit Insights and Marks/Nanda in the pavilions.
No must-attend orals this block. Spend the time at Poster Session 6 in Pav 4: Atticus' Lookbacks paper (P4-#4013), Nanda/Senthooran's Emergent Misalignment (P4-#4802), and Fazl Barez's papers. Good for follow-up conversations.
Key people at the conference, where to find them, and what to discuss.
How to get the most out of three days.
The 10:30 AM block has poster Session 1 in the pavilions and Oral Sessions 1A-1F in rooms 201-204. The SAE cluster posters are in Pav 4 (Temporal SAEs oral, AbsTopK, Amortized Inference, Paulo/Belrose) lets you do a fast sweep of the SAE landscape in one pass. All within walking distance of each other.
The 3:15 PM block is your most important as an attendee. FaithCoT-Bench and Thought Branches are both must-visit posters in Pav 4 during Poster Session 2. Start with FaithCoT-Bench (you need this conversation for your NeurIPS paper), then Thought Branches (meet Nanda, engage with the trajectory-distribution argument), The CoT Computational Graph oral is at 11:06 AM in Room 201 A/B (morning block), so catch that first, then hit the afternoon posters. Also note: "How Do Transformers Learn to Associate Tokens" oral is at 3:27 PM in Room 201 C, and "Temporal SAEs" oral is at 4:03 PM in Room 204 C -- both during the afternoon oral block.
Morning (10:30 AM): Poster Session 3 in the pavilions. Hit RFEval and the GRPO Unlearning poster. Orals 3A-3F run in parallel in rooms 201-204 -- notably GEPA (Potts, Oral 3A, Room 201 A/B, 11:30 AM) and SAEnce (Oral 3F, Room 204 C, 11:30 AM). Save energy for your poster at 3:15 PM.
Poster Session 4 (3:15 PM): You're presenting. Your neighborhood in Pav 4 is stacked -- Markovian Transformers (P4-#4303), In-Context Algebra (P4-#4011), Nanda's planning paper (P4-#4308), and two unlearning papers are all nearby. Between waves of visitors, take 5-minute walks to these posters. David Bau's In-Context Algebra is the highest-priority neighbor visit.
Session 5: More relaxed. Circuit Insights and the Marks/Nanda evaluation-awareness paper. Good for follow-up conversations with people you met on Days 1-2.
Morning orals (10:30 AM): Anchor to the Potts oral on causal interventions (Oral 5D, Room 203 A/B, 11:30 AM). This is your strongest Stanford introduction opportunity on research substance. "Art of Scaling RL Compute" is also in the morning orals (Oral 5C, Room 202 A/B, 11:06 AM). In the afternoon (3:15 PM), Poster Session 6 has Atticus' Lookbacks paper and the Nanda/Senthooran Emergent Misalignment poster in Pav 4.
"We reverse-engineer how language models retrieve the right entity when multiple entities share an attribute. We find a 'mixing mechanism' where the model uses attention to construct hybrid representations that can be read off by later layers. The key finding is that this isn't simple copying -- it's a structured computation where the model mixes entity and attribute information in a specific, recoverable way."
If they ask about implications: "It tells us something about how factual knowledge is actually accessed in-context, which matters for both interpretability and for understanding retrieval failures."
"I'm working on benchmarking CoT faithfulness. The core problem is that existing benchmarks often measure whether the reasoning looks correct rather than whether the model actually used that reasoning. We're building evaluation methods that distinguish plausible-looking chains from genuinely faithful ones."
If they ask how: "We use RL training to create models that demonstrably reason unfaithfully -- they make discriminatory decisions while suppressing the decision-relevant attributes in their chain-of-thought. Then we test whether existing faithfulness metrics can catch this."
Orals are 12-minute talks in rooms 201-204 (second floor). They run 10:30 AM – 12:00 PM and 3:15 PM – 4:45 PM, parallel with poster sessions. Here are the ones worth attending:
Thu 23 · Morning (10:30–12:00)
11:06 · Room 201 A/B · Verifying CoT Reasoning via Computational Graph (Oral 1A)
Thu 23 · Afternoon (3:15–4:45)
3:15 · Room 203 A/B · Beyond Prompt-Induced Lies (Oral 2D)
3:27 · Room 203 A/B · Is it Thinking or Cheating? Detecting Implicit Reward Hacking (Oral 2D)
3:27 · Room 201 C · How Do Transformers Learn to Associate Tokens (Oral 2B)
4:03 · Room 204 C · Temporal Sparse Autoencoders (Oral 2F)
4:15 · Room 203 A/B · What's In My Human Feedback? (Oral 2D)
Fri 24 · Morning (10:30–12:00)
11:30 · Room 201 A/B · GEPA — Christopher Potts (Oral 3A)
11:30 · Room 204 C · Exploratory Causal Inference in SAEnce (Oral 3F)
11:42 · Room 202 A/B · Pre-training under infinite compute — Liang, Hashimoto (Oral 3C)
Fri 24 · Afternoon (3:15–4:45)
4:27 · Room 202 A/B · Visual Symbolic Mechanisms (Oral 4C)
Sat 25 · Morning (10:30–12:00)
11:06 · Room 202 A/B · The Art of Scaling RL Compute for LLMs (Oral 5C)
11:06 · Room 203 A/B · Navigating Latent Space Dynamics (Oral 5D)
11:30 · Room 203 A/B · Addressing Divergent Representations from Causal Interventions — Potts (Oral 5D)
Note: Thu afternoon has the most conflicts. "Transformers Associate Tokens" (201 C) and "Reward Hacking" (203 A/B) overlap at 3:27. Prioritize the one more relevant to your NeurIPS paper (reward hacking), since the Transformers talk is more adjacent to Mixing Mechanisms which you already know well.
Papers flagged: 28 (out of 5,471). Oral talks to attend: 12 across 3 days (see quick reference above). Must-talk-to people: 8. Your poster neighbors worth visiting: 5 within a 2-minute walk.
Pavilion split: Posters are in Pavilion 3 + 4 (ground floor). Orals are in rooms 201-204 (second floor). Almost all your relevant posters are in Pavilion 4.