Dom 🆚 Machine

I'm a security researcher/engineer. This is where I post notes on vulnerability research, exploit development, offensive security, and what I'm learning about AI/ML safety & internals.

Two Routes Through a Transformer

The groundwork I needed before A Mathematical Framework for Transformer Circuits and its walkthrough video made sense. What a one-layer attention-only model actually computes.

·14 min ·transformers, interpretability, transformer-circuits

Sampling from a Transformer

Porting the generation loop into PvML, the strategies to turn logits into a token, and the one-token bug that made a trained model emit nothing but commas.

·9 min ·transformers, arena, samplers

Training a Transformer

Continuation of the Understanding Transformers post, where the weights stop being downloaded and start being learned. A loss, a data pipeline, a trainer, and experiments.

·19 min ·transformers, tinystories, arena

Understanding Transformers

Notes from rebuilding GPT-2 from scratch while working through ARENA's curriculum, module by module, with the shapes and diagrams I needed to follow it.

·24 min ·transformers, gpt-2, arena
↑