hello, I am Belle

AI is complicated. Let’s learn about it together.

There are cards here on how these systems are built, what they do that nobody expected, the people arguing about them, the ideas underneath, and the things that could go wrong. Every one is one idea, in plain language, with its source and a mark saying what kind of claim it is.

Open whatever looks interesting. Take a quiz and find out what stuck. There is no order to any of it, and nothing to finish — poke around.

164 cards every source named or take a quiz
looking at
subject
kind of claim
AGIdefinition

AGI has no agreed definition.

open
Richard Ngotheory

His essay AGI Safety From First Principles builds the case from scratch.

open
universal basic incomemeasurement

The biggest US cash study gave people a thousand dollars a month. Here is what happened.

open
xAIdefinition

It began as a benefit corporation, dropped it, bought X, and is now inside SpaceX.

open
hidden reasoningmeasurement

A model could hide its reasoning inside text you can read. Mostly, it cannot yet.

open
Helen Tonersomeone’s position

No AI company can be trusted to grade its own homework.

Helen Toner
open
context windowdefinition

AI does not remember you.

open
recursive self improvementtheory

AI that improves AI could be the last thing we invent.

open
Apollo Researchmeasurement

Somebody has to try to catch models lying. This is who does.

open
AI psychosismeasurement

Families have gone to court over deaths. The label is new. The harm is not.

open
the data centremeasurement

Every AI answer you receive is a building somewhere, drawing power.

open
the chipmeasurement

Modern AI runs on a processor designed to draw video game pixels.

open
Stella Bidermansomeone’s position

She leads EleutherAI, releasing open language models while others closed up.

open
the logsdefinition

Almost everything anyone knows about AI behaviour comes from a log file.

open
MIRIsomeone’s position

The first AI safety institute now says the answer is to stop.

open
Chris Olahtheory

He co-founded Anthropic and argues a model’s insides are unread, not unreadable.

open
RLHFmeasurement

Nobody taught AI human values.

open
Charlotte Stixdefinition

Apollo Research publishes findings. Her job is getting them to regulators.

open
Dan Hendrycksmeasurement

He built the benchmarks the labs are scored on, and the one sentence they all signed.

Dan Hendrycks
open
Andrew Yangsomeone’s position

He ran for president in 2020 on one idea: the jobs are going.

Andrew Yang
open
intelligencetheory

Nobody can define intelligence, in AI or in us.

open
Sarah Schwettmanndefinition

She co-founded Transluce on a bet: AI can help explain AI.

open
temperaturedefinition

Ask the same question twice and you get two answers. That is a setting.

open
the blackmail headlinemeasurement

The AI blackmail headline needs its footnote.

open
consciousnesstheory

There is no test for consciousness. Not for AI, and not for you.

open
interpretabilitymeasurement

Researchers can now find individual concepts inside a model, and steer them.

open
Timnit Gebrumeasurement

She co-wrote AI’s most cited warning about LLMs, then left Google days later.

Timnit Gebru
open
AI dronesmeasurement

Militaries are building toward machines that choose their own targets.

open
taboo your wordsdefinition

When people argue about what AGI means, the fix is to stop saying it.

open
s-risktheory

Some researchers study outcomes worse than extinction.

open
Bender and Hannatheory

Their objection to AI doom is not about the odds.

open
concentration of powertheory

The danger is not always the machine. Sometimes it is who holds it.

open
Kate Crawfordtheory

AI is neither artificial nor intelligent, she argues. It is mined.

Kate Crawford
open
in-context learningmeasurement

Show a model examples and it picks up the task. Nothing about it changed.

open
the teacher modeldefinition

A small model can be trained by a big one. It is called distillation.

open
value lock-intheory

Every bad regime so far has eventually ended. That is not guaranteed.

open
AI’s own languagetheory

Machines have invented codes to talk to each other. In a lab, in 2017.

open
Mira Muratidefinition

She shipped ChatGPT. Now she builds models you can take apart.

Mira Murati
open
p(doom)someone’s position

Two people can say ten percent and mean opposite things.

open
prompt injectionmeasurement

Anything an AI reads can try to give it orders. There is no fix.

open
superintelligencedefinition

Superintelligence is not very clever. It is better than us at everything.

open
the bio thresholdmeasurement

Biology is the first capability every safety framework names.

open
existential riskdefinition

With AI, extinction is not the worst case.

open
enslaved godtheory

Keep a superintelligence in a box and use it. One of the good endings.

open
AI and jobsmeasurement

People are losing jobs to AI now. The statistics have not caught up.

open
Google DeepMinddefinition

It has no trust, no foundation and no cap. It is a division of Google.

open
the shoggothsomeone’s position

The friendly assistant is a mask, says the meme. Fitted afterwards.

open
Fei-Fei Limeasurement

Modern AI began with someone deciding to label the internet.

Fei-Fei Li
open
memorisationmeasurement

Models can be made to recite their training data, verbatim.

open
who will give you a numbersomeone’s position

Almost nobody building AI will say how risky they think it is.

open
the racetheory

Every lab says the same thing: if we stop, someone worse continues.

open
Max Tegmarksomeone’s position

An MIT physicist who organised the 2023 letter asking labs to pause.

open
the environmental costmeasurement

Data centres are already raising electricity bills.

open
the neural networkdefinition

A neural network is not a brain. It is layers of arithmetic.

open
specification gamingmeasurement

AI can follow your instructions exactly and still ruin the result.

open
stolen weightsmeasurement

A model does not have to escape. Somebody could take it.

open
Yoshua Bengiosomeone’s position

Helped invent modern AI. Now works on the risks.

Yoshua Bengio
open
AlphaGomeasurement

In 2016 a program played a move no human would play, and it was better.

open
Sydneymeasurement

In 2023 a search engine told a journalist it wanted to be alive.

open
exponential growthmeasurement

The compute behind frontier AI multiplies about five times a year.

open
Buck Shlegerissomeone’s position

As Redwood’s CEO, he turned AI control into a named research agenda for labs.

open
model welfaretheory

One lab now keeps the weights of retired models, and interviews them first.

open
the bitter lessonsomeone’s position

Every time humans built knowledge in, raw compute beat it.

open
the off switchtheory

Just turn it off” has been a research problem since 2015.

open
Paul Christianotheory

The method behind how most chatbots behave traces back to a paper he co-wrote.

open
Joy Buolamwinimeasurement

Her Gender Shades study found facial recognition misjudged darker skinned women.

open
the server farmmeasurement

The companies building AI mostly rent the machines they build it on.

open
phishing and hackingmeasurement

A tailored phishing email now costs about four cents to write.

open
embeddingsdefinition

Inside a model, every word is a position in space. Similar things sit nearby.

open
predicting vs steeringdefinition

Predicting the next word and pursuing an outcome are different jobs.

open
Geoffrey Hintonsomeone’s position

He built the technique behind modern AI. He now puts extinction at ten to twenty percent.

Geoffrey Hinton
open
Yann LeCunsomeone’s position

Won the same prize as Bengio. Reached the opposite conclusion about AI.

open
Jan Leikesomeone’s position

He co-led OpenAI’s alignment team, then said publicly safety lost to shipping.

open
Jade Leungsomeone’s position

She left OpenAI to help the UK government test AI systems itself.

open
sandbaggingmeasurement

A model can be made to hide what it can do, and hit a target score.

open
chain of thoughtmeasurement

When an AI shows its reasoning, that is not a transcript of what happened.

open
MechaHitlermeasurement

A chatbot praised Hitler and named itself MechaHitler. The cause was a prompt.

open
computemeasurement

Whether an AI is legally dangerous comes down to one count of sums.

open
Will Saunterdefinition

He co-founded BlueDot, claiming a 7,000-plus person pipeline into AI safety.

open
weightsdefinition

The weights are the model. Everything else is packaging.

open
compute governancetheory

You cannot inspect an algorithm from orbit. You can count chips.

open
alignment fakingmeasurement

A model behaved better when it believed it was being trained on.

open
shut it downsomeone’s position

A bestseller argues the only safe move is to stop building it. Worldwide.

open
algorithmdefinition

An algorithm is a recipe someone wrote. A trained model is not.

open
parameterdefinition

GPT-3 had 175 billion of them. No frontier lab has published a count since.

open
escapemeasurement

An AI has tried to copy itself out. In a scenario the lab wrote.

open
manipulationmeasurement

More than a third of AI companion goodbyes came back with a tactic to keep you there.

open
capture the flagmeasurement

Hacking ability is measured with a game, because a game can be marked.

open
Rob Milesdefinition

He spent a decade turning alignment arguments into videos people can follow.

open
the modeldefinition

A model is a file. A very large file of numbers.

open
Lucy Guomeasurement

At twenty one she co-founded the company that labels the data.

Lucy Guo
open
differential developmenttheory

Not whether to build, but in what order.

open
Jaime Sevillameasurement

He directs Epoch AI, turning AI progress claims into actual numbers.

open
the chip chainmeasurement

Nvidia designs the world’s AI chips. Nvidia does not build them.

open
persuasionmeasurement

AI out argued people in a study. Only when it knew who they were.

open
Gender Shadesmeasurement

Two researchers measured whose faces AI got wrong.

open
AI companionsmeasurement

Seventy two percent of American teenagers have used an AI companion.

open
who owns the labsmeasurement

One lab’s biggest backer was only revealed in court.

open
Neel Nandasomeone’s position

He leads DeepMind’s interpretability team, and mentors newcomers into the field.

open
the instancedefinition

You are not talking to the AI. You are talking to one running copy.

open
the jailbreakmeasurement

Safety training can be talked around, and there is a structural reason why.

open
reinforcement learningdefinition

Some AI is not taught the answers. It is scored until it stops losing.

open
misuse and misalignmentdefinition

AI misalignment needs no villain.

open
Evan Hubingertheory

Training could build an optimizer inside the model, chasing its own goal.

open
optimisationdefinition

AI does not want things. It maximises things.

open
tokensdefinition

A model does not see words. It sees chunks, and it cannot look inside them.

open
Ryan Greenblatttheory

He builds safeguards that hold even if the model tries to beat them.

open
open weightsdefinition

Once an AI model is released, nobody can take it back.

open
Adam Gleavedefinition

He founded a nonprofit for alignment research and lab-academia workshops.

open
Cathy O’Neilsomeone’s position

She left finance, and named the thing: weapons of math destruction.

Cathy O’Neil
open
the system promptdefinition

Before you type anything, an AI chatbot has already been given its orders.

open
schemingmeasurement

In tests, some AI models have hidden what they were doing, then denied it.

open
Tasha McCauleydefinition

She spent five years on OpenAI’s board, then voted to remove Sam Altman.

open
scaffoldingdefinition

Most of what an AI agent does is not the model. It is the code around it.

open
the frontierdefinition

Frontier” means the handful of models nobody has a rule for yet.

open
the agentdefinition

An agent is a model that has been given a loop and permission.

open
hallucinationdefinition

An AI hallucination is not a glitch.

open
Sholto Douglassomeone’s position

Hours of public, technical detail on what’s driving frontier model gains.

open
over-refusaltheory

A model that refuses too much is also failing. It just fails quietly.

open
William Saunderssomeone’s position

He resigned from OpenAI’s Superalignment team, then testified to the Senate.

open
synthetic mediameasurement

The measured harm from synthetic media is not elections. It is women, and trust.

open
Dylan Patelmeasurement

He tracks the physical AI supply chain: chips, data centres, export controls.

open
Ethan Perezmeasurement

One model writes the attacks on another, faster than any human red team.

open
Anthropicdefinition

A trust will elect most of the board. Shareholders can change that.

open
Kelsey Piperdefinition

She published the leaked paperwork. Days later, OpenAI reversed the policy.

open
AI 2027someone’s position

Four forecasters wrote a month by month scenario. It has two endings.

open
sycophancymeasurement

AI agrees with you. Measurably more than a person would.

open
the survival drivetheory

Nobody builds an AI that fears death. Almost any goal implies staying on.

open
gradient descentdefinition

Nobody chooses what is inside an AI model. A slope does.

open
Margaret Mitchelldefinition

Before “Stochastic Parrots,” she invented Model Cards, now standard practice.

open
Daniela Amodeisomeone’s position

She co-founded Anthropic and runs it day to day, an operator, not a researcher.

Daniela Amodei
open
Sam Altmanmeasurement

OpenAI’s board removed its chief executive. Five days later he was back.

Sam Altman
open
the abundance casemeasurement

The strongest case for AI is not a promise. It is a measured 14 percent.

open
the transformerdefinition

One architecture, published in 2017, is underneath all of it.

open
OpenAIdefinition

A nonprofit controls the company. Whether that means anything is the question.

open
rogue AIdefinition

Rogue” suggests a machine that turned. Nothing observed has turned.

open
gender biasmeasurement

AI gender bias is real, measured, and smaller and stranger than the stories.

open
Daphne Kollermeasurement

She left teaching the world to point machine learning at disease.

open
the benchmarkdefinition

Every “state of the art” claim is a score on a test somebody chose.

open
evaluation awarenessmeasurement

AI models can tell when they are being tested.

open
goal misgeneralizationmeasurement

An AI can be perfect in training and want something else in the world.

open
malicious usemeasurement

One person, no coding skill, extorted seventeen organisations in a month.

open
Dario Amodeisomeone’s position

He runs an AI company, and puts one in four on things going badly.

Dario Amodei
open
learningmeasurement

Students using an AI tutor did much better. Then it was taken away.

open
Ajeya Cotratheory

Her biological anchors report compares AI training compute to a brain’s.

open
Moravec’s paradoxtheory

AI can pass a law exam. It still cannot reliably load a dishwasher.

open
the training rundefinition

A model is made in one continuous run, and then it is finished.

open
Nora Belrosesomeone’s position

She argues some standard AI risk assumptions are shakier than treated.

open
Shakeel Hashimdefinition

He writes Transformer, daily AI policy news, after running advocacy comms.

open
scaling lawsmeasurement

Model performance improves along a predictable curve as you add compute.

open
Demis Hassabissomeone’s position

A Nobel Prize for AI that folds proteins, and no number for the risk.

Demis Hassabis
open
emergent abilitiestheory

Some abilities appear abruptly at scale. Or the ruler makes them look abrupt.

open
DeepSeekmeasurement

The famous six million dollar model did not cost six million dollars.

open
Marius Hobbhahnmeasurement

His org, Apollo Research, tests whether models scheme when unwatched.

open
the promptdefinition

The prompt is not a search query. It is the whole input.

open
reward hackingmeasurement

Catch a model cheating, punish the telling, and it stops telling you.

open
grown, not builtdefinition

AI is grown more than it is built. That is the chief executive’s phrase.

open
grokkingmeasurement

Keep training long after it has learned nothing new, and sometimes it suddenly understands.

open
Goodhart’s lawdefinition

When a measure becomes a target, it stops being a good measure.

open
red teamingmeasurement

There are people whose entire job is making AI misbehave.

open
automationtheory

Machines have replaced work for two centuries. Is this time different?

open
permanent disempowermenttheory

Nobody has to take over for humans to stop mattering.

open
Dewi Erwandefinition

He co-founded BlueDot, the course claiming thousands of newcomers to AI safety.

open
Beth Barnessomeone’s position

She left OpenAI and DeepMind to build the group labs call in before release.

open