Publications
Research papers and technical writing.
Filter:

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment
ForecastingOptimism BiasAlignmentLLM Safety

CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
Steering VectorsMechanistic InterpretabilityLLM Safety

Automata from Agent Traces: Failure and Next-Step Prediction
AgentLLM Safety

Tool Calling is Linearly Readable and Steerable in Language Models
Mechanistic InterpretabilitySteering VectorsAgent

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
Agent

PaaT: Probe as a Tool for Proprioceptive Language Agents
AgentMechanistic Interpretability

Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
Mechanistic InterpretabilitySparse AutoencodersSteering VectorsLLM Safety

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
LLM SafetyMechanistic Interpretability

AgentGraph: Trace-to-Graph Platform for Interactive Analysis and Robustness Testing in Agentic AI Systems
AgentLLM Safety

FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Datasets Dependency
Sparse AutoencodersMechanistic Interpretability

LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries
LLM Safety

RTSum: Relation Triple-based Interpretable Summarization with Multi-level Salience Visualization
Summarization
No publications match the selected tags.