WriteSAE

WriteSAE uses a sparse autoencoder to study what a recurrent language model stores in memory and how it affects predictions. This recurrent state is a matrix that carries information between tokens. We reconstruct saved states from a few active features, each shaped like a native memory write.

  • Inspect which inputs activate a feature to study what information the state carries.
  • Use the experiments to ablate or modify features in the state or its memory writes, then measure the effect on later predictions. Memory edits changed token scores as predicted (median R² = 0.98 on Qwen3.5-0.8B, layer 9, head 4). The language model weights stay frozen.

Load a trained autoencoder

python -m pip install huggingface_hub torch
hf download JackYoung27/writesae-ckpts --local-dir writesae --include 'core/*' 'LOAD_EXAMPLE.py' 'writesae/qwen0p8b/L9_H4/*'
cd writesae
python LOAD_EXAMPLE.py

This example loads the Qwen3.5-0.8B layer 9, head 4 autoencoder on a CPU and checks the output shapes.

Files

  • writesae/: seven trained autoencoder packs, each with weights and settings.
  • flat_baseline/: comparison checkpoints.
  • core/: autoencoder and training code.
  • experiments/ and scripts/: code for testing memory edits.
  • results/: saved experiment outputs.
  • manifest.json: artifact metadata and checksums.
  • Paper.

The loading example checks reconstruction. It does not run a memory-editing experiment or load the base language model.

MIT license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JackYoung27/writesae-ckpts

Finetuned
(367)
this model