Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
"Matryoshka attribution: Learning to attribute language model outputs to representations and weights" What if you could trace a model behavior back to the smallest set of internal components actually responsible for it? Matryoshka Attribution here basically learns one ranked mask over internal components across many sparsity levels at once, letting it isolate compact causal circuits efficiently. This approach reaches SoTA circuit localization, and can even attribute behaviors to weight…
