Personal Conference Guide

ICLR 2026 · Rio de Janeiro

April 23–27  ·  5,471 papers reviewed  ·  28 flagged for you

★ Your Poster

Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context

Yoav Gur-Arieh, Mor Geva, Atticus Geiger  ·  Session 4 · Pavilion 4 · P4-#4306
Thursday April 24, 3:15 PM – 5:45 PM

Oral
Must-see (CoT / faithfulness)
Mech Interp / SAEs
Unlearning
Safety / alignment

Wednesday, April 23

Heaviest day — the mech interp and SAE cluster lands here. Two sessions back-to-back.

10:30 – 13:00 Session 1 · Pavilion 3 + 4
Temporal Sparse Autoencoders ORAL
Bhalla et al.
Oral 2F · Room 204 C · Thu 4:03 PM (also poster: Pav 4, P4-#4306)
Extends SAEs to leverage sequential token structure rather than treating activations as i.i.d. Could change how we think about feature extraction in practice. Worth catching the talk if you can make the early slot.
How Do Transformers Learn to Associate Tokens ORAL
Im et al.
Oral 2B · Room 201 C · Thu 3:27 PM (also poster: Pav 4, P4-#4006)
Mechanistic account of token association via gradient leading terms. Directly relevant to how binding works in transformers, which is your Mixing Mechanisms territory.
Is it Thinking or Cheating? Detecting Implicit Reward Hacking in CoT ORAL
 
Oral 2D · Room 203 A/B · Thu 3:27 PM
Reward hacking in chain-of-thought is exactly the failure mode your NeurIPS work targets. Useful framing to compare against.
Toward Faithful RAG with Sparse Autoencoders INTERP
Xiong et al.
Pav 4, P4-#3811
Uses SAEs to mechanistically analyze RAG faithfulness. Intersection of interp + faithfulness that overlaps with both your research threads.
AbsTopK: Rethinking Sparse Autoencoders for Bidirectional Features INTERP
 
Pav 4, P4-#3911
SAE architecture variant. Relevant to your Goodfire work and general SAE methodology.
The Price of Amortized Inference in Sparse Autoencoders INTERP
 
Pav 4, P4-#4008
Addresses polysemy in SAE features. Fundamental limitation that Goodfire-adjacent work needs to grapple with.
Sparse Autoencoders Trained on Same Data Learn Different Features INTERP
Paulo, Belrose
Pav 4, P4-#4004
Reproducibility concern for the entire SAE-based interp paradigm. Important methodological result.
Tracking Equivalent Mechanistic Interpretations Across Networks INTERP
Sun, Toneva
Pav 3, P3-#1011
Cross-network mechanistic comparison. Different pavilion from the SAE cluster, so plan your route.
15:15 – 17:45 Session 2 · Pavilion 3 + 4
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of CoT Reasoning MUST SEE
Shen et al.
Pav 4, P4-#4006
This is the Shen et al. paper you've been critiquing in your NeurIPS submission. You need to talk to the authors. Understand their response to the conflation-of-plausibility-with-faithfulness critique. This conversation could shape your paper's related work framing significantly.
Thought Branches: Interpreting LLM Reasoning Requires Resampling MUST SEE
Macar, Bogdan, Senthooran Rajamanoharan, Neel Nanda
Pav 4, P4-#4014
Argues that single CoT traces are insufficient for interpreting reasoning -- you need to sample the distribution of trajectories. This is a methodological challenge to any single-trace faithfulness benchmark, including yours. Critical to engage with. Also your first shot at meeting Neel Nanda.
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations MUST SEE
Ferreira, Aziz, Titov
Pav 4, P4-#3911
Demonstrates that CoT explanations can be reward-hacked and proposes causal attribution as a fix. Directly parallels your GRPO experiments showing models suppress decision-relevant attributes in CoT. Compare approaches.
Verifying Chain-of-Thought Reasoning via Its Computational Graph ORAL
Zhao et al.
Oral 1A · Room 201 A/B · Thu 11:06 AM (also poster: Pav 4, P4-#4115)
Formalizes CoT verification through computational graph structure. A different angle on the faithfulness problem that complements your probability-based approach.
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences INTERP
Minder et al., Neel Nanda
Pav 4, P4-#4005
Model diffing via activations. Second Nanda paper in this session. If you miss him at Thought Branches, catch him here.

Thursday, April 24

Your poster day. Session 3 in the morning, then you present in Session 4.

10:30 – 13:00 Session 3 · Pavilion 3 + 4
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Intervention MUST SEE
Han et al.
Pav 4, P4-#3902
Formal framework for reasoning faithfulness using counterfactual interventions. Another faithfulness benchmark to position against in your NeurIPS paper. Understand how they define and operationalize "faithful" compared to your approach.
Reinforcement Unlearning via GRPO UNLEARN
Zaradoukas et al.
Pav 4, P4-#3905
Uses GRPO for unlearning, which is the same RL method you've been using in your NeurIPS experiments. Different application (unlearning vs. inducing unfaithful CoT), but the methodology overlap is worth exploring.
From Data Statistics to Feature Geometry: How Correlations Shape Superposition INTERP
Prieto et al.
Pav 4, P4-#4413
Theoretical work on superposition. Relevant to how features get packed and read out in the binding mechanisms you study.
Visual Symbolic Mechanisms ORAL
 
Oral 4C · Room 202 A/B · Fri 4:27 PM
Emergent symbol processing in vision. Cross-modality parallel to the symbolic binding mechanisms in your language model work.
15:15 – 17:45 Session 4 · Pavilion 3 + 4 YOU PRESENT
Markovian Transformers for Informative Language Modeling MUST SEE
Viteri et al.
Pav 4, P4-#4303
Introduces a Markovian constraint that forces CoT to faithfully reflect internal decisions -- a "reasoning bottleneck." This is an architectural approach to the faithfulness problem, complementary to your evaluation-based approach. Right next to your poster. Swing by between visitors.
In-Context Algebra INTERP
Eric Todd, David Bau et al.
Pav 4, P4-#4011
Mechanisms for variable binding in transformers. Extremely close to your Mixing Mechanisms work. Compare their "algebra" framing with your entity binding circuits. Good chance to connect with David Bau.
Exploratory Causal Inference in SAEnce ORAL
Mencattini, Cadei, Locatello
Oral 3F · Room 204 C · Fri 11:30 AM (not Pav 3 — that was the old poster ref)
Causal inference over SAE features. Connects your interp background to causal methods. In Pavilion 3 though, so you'd need to step away from your poster.
What's the plan? Metrics for implicit planning in LLMs SAFETY
Neel Nanda et al.
Pav 4, P4-#4308
Nanda paper two posters away from yours. If you didn't connect on Day 1, this is a natural opportunity.
CLUE: Conflict-guided Localization for LLM Unlearning UNLEARN
 
Pav 4, P4-#4310
Circuit discovery for unlearning. Right near your poster. Bridges your interp and unlearning interests.
Unlearning Isn't Invisible: Detecting Unlearning Traces from Model Outputs UNLEARN
Chen et al.
Pav 4, P4-#3910
Evaluates whether unlearning is detectable. Relevant if your unlearning research direction continues.
Safety Subspaces are Not Linearly Distinct SAFETY
Ponkshe et al.
Pav 4, P4-#4204
Challenges linear representation hypothesis for safety features. Important methodological caution for interp-based safety work.
The Lattice Representation Hypothesis INTERP
Bo Xiong
Pav 4, P4-#4405
Proposes lattice structure in LLM embeddings. Theoretical interp work worth skimming if traffic at your poster is slow.
Pre-training under infinite compute ORAL
Percy Liang, Tatsunori Hashimoto et al.
Oral 3C · Room 202 A/B · Fri 11:42 AM
Stanford networking target. The oral is in Pav 3 and conflicts with your poster, but the names are worth knowing for hallway conversations.
GEPA ORAL
Christopher Potts
Oral 3A · Room 201 A/B · Fri 11:30 AM
Potts oral. Same conflict, but worth noting for a hallway encounter.

Friday, April 25

Post-poster day. Free to roam. Safety/alignment cluster lands here, plus Potts' causal interventions oral.

10:30 – 13:00 Session 5 · Pavilion 3 + 4
Circuit Insights: Towards Interpretability Beyond Activations INTERP
Golimblevskaia et al.
Pav 4, P4-#4017
Pushes interp past activation-level analysis toward circuit-level understanding. Methodologically aligned with your approach in Mixing Mechanisms.
Steering Evaluation-Aware LMs To Act Like They Are Deployed SAFETY
Hua, Qin, Samuel Marks, Neel Nanda
Pav 4, P4-#4018
Models that behave differently when they detect evaluation. Connects to faithfulness (is the model's behavior the same as its "reasoning"?). Marks + Nanda together here.
Evidence for Limited Metacognition in LLMs SAFETY
Ackerman
Pav 3, P3-#603
Do models know what they know? Relevant to the faithfulness question -- if metacognition is limited, CoT faithfulness has a ceiling.
Jailbreak Transferability Emerges from Shared Representations SAFETY
Angell et al.
Pav 4, P4-#5116
Mechanistic analysis of jailbreaks. Less central to your work but interesting interp-meets-safety angle.
15:15 – 17:45 Session 6 · Pavilion 3 + 4
Addressing Divergent Representations from Causal Interventions ORAL
Christopher Potts
Oral 5D · Room 203 A/B · Sat 11:30 AM
Potts' second oral, and this one is directly relevant to your methods. Causal intervention methodology is central to your interp work. Strong reason to attend and introduce yourself afterwards. Stanford connection.
Emergent Misalignment is Easy, Narrow Misalignment is Hard SAFETY
Soligo et al., Neel Nanda, Senthooran Rajamanoharan
Pav 4, P4-#4802
Last Nanda paper of the conference. If you still haven't had a real conversation with him by now, this is your final window.
Language Models Use Lookbacks to Track Beliefs SAFETY
Atticus Geiger et al.
Pav 4, P4-#4013
Atticus' other paper. You obviously know this one, but stop by to support your co-author and meet his other collaborators.
Navigating the Latent Space Dynamics of Language Models ORAL
 
Oral 5D · Room 203 A/B · Sat 11:06 AM
Latent space dynamics. Theoretical interp that could inform how you think about information flow in binding circuits.
The Art of Scaling RL Compute for LLMs ORAL
 
Oral 5C · Room 202 A/B · Sat 11:06 AM
Relevant to your GRPO training work. Understanding RL compute scaling helps calibrate the NeurIPS experiments.

Oral Sessions

Rooms 201-204 (second floor). Each talk is ~12 min. Orals run parallel to poster sessions.

Thu 23 · 10:30 – 12:00 Oral Block 1
Verifying Chain-of-Thought Reasoning via Its Computational Graph ORAL 1A
Zhao et al.
11:06 AM · Room 201 A/B
Formalizes CoT verification through computational graph structure. Different angle on faithfulness that complements your probability-based approach.

Nothing else in this block -- head to Poster Session 1 in Pav 4 for the SAE cluster (Temporal SAEs poster, AbsTopK, Amortized Inference, Paulo/Belrose) after this talk.

Thu 23 · 3:15 – 4:45 Oral Block 2 ⭐ Busiest slot
Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts ORAL 2D
 
3:15 PM · Room 203 A/B
LLM deception without adversarial prompting. Sets the stage for the reward hacking talk right after it in the same room.
Is it Thinking or Cheating? Detecting Implicit Reward Hacking in CoT ORAL 2D
 
3:27 PM · Room 203 A/B
Top priority. Reward hacking in CoT is exactly the failure mode your NeurIPS work targets. Stay in 203 A/B after the previous talk.
How Do Transformers Learn to Associate Tokens ORAL 2B
Im et al.
3:27 PM · Room 201 C ⚠️ conflicts with above
Mechanistic token association via gradient leading terms. Relevant to binding in Mixing Mechanisms, but conflicts with the reward hacking talk. Catch the poster instead.
Temporal Sparse Autoencoders ORAL 2F
Bhalla et al.
4:03 PM · Room 204 C
SAEs leveraging sequential token structure. Walk from 203 A/B after the 2D talks finish (~3:50). You have time.
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data ORAL 2D
 
4:15 PM · Room 203 A/B
Interpretable preference descriptions. Connects to RLHF and understanding what reward signals actually capture.

Recommended route: Start in 203 A/B for the 2D session (3:15-3:50ish). Dash to 204 C for Temporal SAEs at 4:03. If you want, return to 203 A/B for the 4:15 talk. Poster Session 2 runs in parallel in the pavilions -- hit FaithCoT-Bench and Thought Branches (Nanda) before or after the oral block.

Fri 24 · 10:30 – 12:00 Oral Block 3 · Stanford morning
GEPA: Reflective Prompt Evolution Can Outperform RL ORAL 3A
Agrawal, Tan, ..., Christopher Potts, ..., Khattab
11:30 AM · Room 201 A/B
Potts co-authored. Prompt optimization outperforming GRPO -- interesting contrast since you use GRPO in your NeurIPS experiments. Good excuse to introduce yourself to Potts after.
Exploratory Causal Inference in SAEnce ORAL 3F
Mencattini, Cadei, Locatello
11:30 AM · Room 204 C ⚠️ conflicts with GEPA
Causal inference over SAE features. Relevant but conflicts with GEPA. Prioritize GEPA for the Potts connection; catch this one at the poster.
Pre-training under Infinite Compute ORAL 3C
Percy Liang, Tatsunori Hashimoto et al.
11:42 AM · Room 202 A/B
Liang + Hashimoto. 12 minutes after GEPA ends in a different room, so you can make both. Stanford double feature.

Recommended route: Go to 201 A/B for GEPA at 11:30. When it ends (~11:42), walk to 202 A/B for the Liang/Hashimoto talk. Two Stanford networking opportunities back to back. Poster Session 3 runs in the pavilions (RFEval, GRPO Unlearning).

Fri 24 · 3:15 – 4:45 Oral Block 4 · You're presenting YOUR POSTER
Visual Symbolic Mechanisms ORAL 4C
 
4:27 PM · Room 202 A/B
Emergent symbol processing in vision. Cross-modal parallel to your binding work. Late enough that you could leave your poster briefly if traffic is light, but your poster comes first.

Only one relevant oral this block -- you'll mostly be at your poster (Pav 4, P4-#4306). Your poster neighbors (In-Context Algebra, Markovian Transformers, Nanda's planning paper, unlearning papers) are the priority.

Sat 25 · 10:30 – 12:00 Oral Block 5 · Potts causal interventions
The Art of Scaling Reinforcement Learning Compute for LLMs ORAL 5C
Devvrit, Madaan, Tiwari, Bansal, Duvvuri, Zaheer, Dhillon, Brandfonbrener, Agarwal
11:06 AM · Room 202 A/B
RL compute scaling. Relevant to your GRPO training experiments and understanding what's feasible at different compute budgets.
Navigating the Latent Space Dynamics of Neural Models ORAL 5D
 
11:06 AM · Room 203 A/B ⚠️ conflicts with above
Theoretical interp on latent dynamics. Conflicts with RL scaling; the RL talk is probably more useful for your NeurIPS paper.
Addressing Divergent Representations from Causal Interventions on Neural Networks KEY TALK
Satchel Grant, Simon Jerome Han, Alexa R. Tartaglini, Christopher Potts
11:30 AM · Room 203 A/B
Your most important oral of the conference. Asks whether causal interventions create OOD representations that undermine interp claims -- directly challenges the methodology behind activation patching and DAS, which are central to your Mixing Mechanisms work. Approach Potts after this talk. If you saw GEPA on Friday, this is your second touchpoint.

Recommended route: Start in 202 A/B for RL Scaling at 11:06. At 11:18 (when it ends), move to 203 A/B for the Potts talk at 11:30. This is your best networking moment of the conference. Poster Session 5 has Circuit Insights and Marks/Nanda in the pavilions.

Sat 25 · 3:15 – 4:45 Oral Block 6

No must-attend orals this block. Spend the time at Poster Session 6 in Pav 4: Atticus' Lookbacks paper (P4-#4013), Nanda/Senthooran's Emergent Misalignment (P4-#4802), and Fazl Barez's papers. Good for follow-up conversations.

Who to Talk To

Key people at the conference, where to find them, and what to discuss.

Neel Nanda
Google DeepMind · 5 papers across Sessions 2, 4, 5, 6
Best chance: Session 2 (Apr 23, 15:15) at "Thought Branches" (P4-#4014) or "Narrow Finetuning" (P4-#4005), both in Pav 4.
Backup: Session 4 — "What's the Plan?" is at P4-#4308, two spots from your poster.
Also at: Session 5 with Samuel Marks, Session 6 with "Emergent Misalignment."
Talk about: Your CoT faithfulness work and how it relates to his "Thought Branches" argument about needing trajectory distributions rather than single traces. Ask how he sees the SAE interpretability agenda connecting to reasoning faithfulness. Given your Goodfire connection, there's natural common ground on feature-level analysis.
Christopher Potts
Stanford NLP · 2 Orals (Sessions 4, 6)
Orals: "GEPA" (Fri 11:30 AM, Room 201 A/B) and "Addressing Divergent Representations from Causal Interventions" (Sat 11:30 AM, Room 203 A/B).
Priority: The causal interventions oral (Sat 11:30 AM, Room 203 A/B). Attend the talk and approach afterwards.
Talk about: Causal intervention methodology in interp, and how your entity binding work (Mixing Mechanisms) uses similar tools. This is your strongest Stanford connection on research substance. Mention your interest in visiting researcher pathways if the conversation goes well.
Percy Liang
Stanford CRFM · 8 papers including Oral in Session 4
Oral: "Pre-training under infinite compute" (Fri 11:42 AM, Room 202 A/B) — conflicts with your poster session but not your poster (your poster is the afternoon).
Multiple posters across Sessions 1-3.
Talk about: HELM and evaluation methodology — your NeurIPS faithfulness benchmarking work is in the same spirit as his systematic evaluation agenda. Could be a less crowded conversation than at the oral.
Tatsunori Hashimoto
Stanford · 6 papers
Co-author on the Liang oral. Multiple posters across Sessions 1-3.
Talk about: RLHF and reward hacking. Your GRPO experiments showing models can learn to suppress information in CoT is directly relevant to his work on alignment training failures.
David Bau
Northeastern · 3 papers (Sessions 1, 4, 6)
Priority: "In-Context Algebra" (Session 4, P4-#4011) — same session as your poster, both in Pav 4.
Talk about: His "In-Context Algebra" and your "Mixing Mechanisms" are studying the same phenomenon (variable binding in transformers) from different angles. This is a natural collaboration or at least a mutual-citation conversation. Walk over during a lull at your poster.
Samuel Marks
Anthropic · Session 5 with Nanda
At: "Steering Evaluation-Aware LMs" (Session 5, P4-#4018)
Talk about: Evaluation-aware behavior connects to CoT faithfulness -- if models change behavior under evaluation, their CoT faithfulness may also be evaluation-contingent.
Shen et al. (FaithCoT-Bench)
Session 2, Pav 4
At: P4-#4006, Session 2 (Apr 23, 15:15)
You are critiquing their benchmark in your NeurIPS paper. This conversation is essential. Be constructive: explain where you think plausibility and faithfulness come apart, and what your approach does differently. Getting their perspective will strengthen your paper.
Noah Goodman
Stanford · 1 paper, Session 6, Pav 4
Session 6 poster in Pav 4.
Stanford target. Probabilistic models of language and cognition. Could be interested in the probability-based faithfulness measurement in your NeurIPS work.

Conference Strategy

How to get the most out of three days.

Day 1 (Apr 23): Reconnaissance

The 10:30 AM block has poster Session 1 in the pavilions and Oral Sessions 1A-1F in rooms 201-204. The SAE cluster posters are in Pav 4 (Temporal SAEs oral, AbsTopK, Amortized Inference, Paulo/Belrose) lets you do a fast sweep of the SAE landscape in one pass. All within walking distance of each other.

The 3:15 PM block is your most important as an attendee. FaithCoT-Bench and Thought Branches are both must-visit posters in Pav 4 during Poster Session 2. Start with FaithCoT-Bench (you need this conversation for your NeurIPS paper), then Thought Branches (meet Nanda, engage with the trajectory-distribution argument), The CoT Computational Graph oral is at 11:06 AM in Room 201 A/B (morning block), so catch that first, then hit the afternoon posters. Also note: "How Do Transformers Learn to Associate Tokens" oral is at 3:27 PM in Room 201 C, and "Temporal SAEs" oral is at 4:03 PM in Room 204 C -- both during the afternoon oral block.

Day 2 (Apr 24): Your Day

Morning (10:30 AM): Poster Session 3 in the pavilions. Hit RFEval and the GRPO Unlearning poster. Orals 3A-3F run in parallel in rooms 201-204 -- notably GEPA (Potts, Oral 3A, Room 201 A/B, 11:30 AM) and SAEnce (Oral 3F, Room 204 C, 11:30 AM). Save energy for your poster at 3:15 PM.

Poster Session 4 (3:15 PM): You're presenting. Your neighborhood in Pav 4 is stacked -- Markovian Transformers (P4-#4303), In-Context Algebra (P4-#4011), Nanda's planning paper (P4-#4308), and two unlearning papers are all nearby. Between waves of visitors, take 5-minute walks to these posters. David Bau's In-Context Algebra is the highest-priority neighbor visit.

Day 3 (Apr 25): Networking + Safety

Session 5: More relaxed. Circuit Insights and the Marks/Nanda evaluation-awareness paper. Good for follow-up conversations with people you met on Days 1-2.

Morning orals (10:30 AM): Anchor to the Potts oral on causal interventions (Oral 5D, Room 203 A/B, 11:30 AM). This is your strongest Stanford introduction opportunity on research substance. "Art of Scaling RL Compute" is also in the morning orals (Oral 5C, Room 202 A/B, 11:06 AM). In the afternoon (3:15 PM), Poster Session 6 has Atticus' Lookbacks paper and the Nanda/Senthooran Emergent Misalignment poster in Pav 4.

Poster Pitch (30-second version)

"We reverse-engineer how language models retrieve the right entity when multiple entities share an attribute. We find a 'mixing mechanism' where the model uses attention to construct hybrid representations that can be read off by later layers. The key finding is that this isn't simple copying -- it's a structured computation where the model mixes entity and attribute information in a specific, recoverable way."

If they ask about implications: "It tells us something about how factual knowledge is actually accessed in-context, which matters for both interpretability and for understanding retrieval failures."

NeurIPS Framing (for conversations)

"I'm working on benchmarking CoT faithfulness. The core problem is that existing benchmarks often measure whether the reasoning looks correct rather than whether the model actually used that reasoning. We're building evaluation methods that distinguish plausible-looking chains from genuinely faithful ones."

If they ask how: "We use RL training to create models that demonstrably reason unfaithfully -- they make discriminatory decisions while suppressing the decision-relevant attributes in their chain-of-thought. Then we test whether existing faithfulness metrics can catch this."

Oral Talks Quick Reference

Orals are 12-minute talks in rooms 201-204 (second floor). They run 10:30 AM – 12:00 PM and 3:15 PM – 4:45 PM, parallel with poster sessions. Here are the ones worth attending:

Thu 23 · Morning (10:30–12:00)
11:06 · Room 201 A/B · Verifying CoT Reasoning via Computational Graph (Oral 1A)

Thu 23 · Afternoon (3:15–4:45)
3:15 · Room 203 A/B · Beyond Prompt-Induced Lies (Oral 2D)
3:27 · Room 203 A/B · Is it Thinking or Cheating? Detecting Implicit Reward Hacking (Oral 2D)
3:27 · Room 201 C · How Do Transformers Learn to Associate Tokens (Oral 2B)
4:03 · Room 204 C · Temporal Sparse Autoencoders (Oral 2F)
4:15 · Room 203 A/B · What's In My Human Feedback? (Oral 2D)

Fri 24 · Morning (10:30–12:00)
11:30 · Room 201 A/B · GEPA — Christopher Potts (Oral 3A)
11:30 · Room 204 C · Exploratory Causal Inference in SAEnce (Oral 3F)
11:42 · Room 202 A/B · Pre-training under infinite compute — Liang, Hashimoto (Oral 3C)

Fri 24 · Afternoon (3:15–4:45)
4:27 · Room 202 A/B · Visual Symbolic Mechanisms (Oral 4C)

Sat 25 · Morning (10:30–12:00)
11:06 · Room 202 A/B · The Art of Scaling RL Compute for LLMs (Oral 5C)
11:06 · Room 203 A/B · Navigating Latent Space Dynamics (Oral 5D)
11:30 · Room 203 A/B · Addressing Divergent Representations from Causal Interventions — Potts (Oral 5D)

Note: Thu afternoon has the most conflicts. "Transformers Associate Tokens" (201 C) and "Reward Hacking" (203 A/B) overlap at 3:27. Prioritize the one more relevant to your NeurIPS paper (reward hacking), since the Transformers talk is more adjacent to Mixing Mechanisms which you already know well.

Quick Stats

Papers flagged: 28 (out of 5,471). Oral talks to attend: 12 across 3 days (see quick reference above). Must-talk-to people: 8. Your poster neighbors worth visiting: 5 within a 2-minute walk.

Pavilion split: Posters are in Pavilion 3 + 4 (ground floor). Orals are in rooms 201-204 (second floor). Almost all your relevant posters are in Pavilion 4.