“The researchers looked inside one of Anthropic’s A.I. models — Claude 3 Sonnet, a version of the company’s Claude 3 language model — and used a technique known as “dictionary learning” to uncover patterns in how combinations of neurons, the mathematical units inside the A.I. model, were activated when Claude was prompted to talk about certain topics.”
– Kevin Roose
A research team at Anthropic says they have identified features within their AI model that provide insights into how large language models work. Anthropic is one of the leading AI developers and made the finding on their proprietary model, Claude-Sonnet.
A.I.’s Black Boxes Just Got a Little Less Mysterious | THE NEW YORK TIMES | May 21, 2024 | by Kevin Roose