May 28, 2024
One of the biggest challenges facing artificial intelligence companies is that they don’t know everything about their algorithms. This so-called black box problem is exacerbated by the fact that deep learning models do precisely that — they learn. And when they learn they change. They take in enormous troves of data, detect patterns, and spit something out: How a sentence should read, what an image should look like, how a voice should sound.
But now researchers at Anthropic, the AI startup that makes the chatbot Claude, claim they’ve had a breakthrough in understanding their own model. In a blog post, Anthropic researchers disclosed that they’ve found 10 million “features” of their Claude 3 Sonnet language model, with certain patterns that pop up when a user inputs something it recognizes. They’ve been able to map features that are close to one another: One for the Golden Gate Bridge, for example, is close to another for Alcatraz Island, the Golden State Warrior, California Governor Gavin Newsom, and the Alfred Hitchcock film Vertigo — set in San Francisco. Knowing about these features allows Anthropic to turn them on or off, manipulating the model to break out of its typical mold.
More For You
Humanitarian aid isn’t just charity—it’s the world’s safety net against disease, conflict, and instability. David Miliband joins Ian Bremmer to unpack what America’s retreat from global aid means for the world and what comes next.
Most Popular
Think you know what's going on around the world? Here's your chance to prove it.
Iran-linked cyberattacks on US water systems show how a war thousands of miles away can reach Americans through something as basic as the tap.
As America celebrates 250 years, Bank of America is proud to honor our shared history. Discover our commitment.
© 2025 GZERO Media. All Rights Reserved | A Eurasia Group media company.
