AGI definition
AGI has no agreed definition.
open →
Richard Ngo theory
His essay AGI Safety From First Principles builds the case from scratch .
open →
universal basic income measurement
The biggest US cash study gave people a thousand dollars a month . Here is what happened.
open →
xAI definition
It began as a benefit corporation, dropped it, bought X, and is now inside SpaceX .
open →
hidden reasoning measurement
A model could hide its reasoning inside text you can read . Mostly, it cannot yet.
open →
Helen Toner someone’s position
No AI company can be trusted to grade its own homework .
open →
context window definition
AI does not remember you.
open →
recursive self improvement theory
AI that improves AI could be the last thing we invent .
open →
Apollo Research measurement
Somebody has to try to catch models lying . This is who does.
open →
AI psychosis measurement
Families have gone to court over deaths. The label is new. The harm is not.
open →
the data centre measurement
Every AI answer you receive is a building somewhere, drawing power .
open →
the chip measurement
Modern AI runs on a processor designed to draw video game pixels .
open →
Stella Biderman someone’s position
She leads EleutherAI, releasing open language models while others closed up.
open →
the logs definition
Almost everything anyone knows about AI behaviour comes from a log file .
open →
MIRI someone’s position
The first AI safety institute now says the answer is to stop .
open →
Chris Olah theory
He co-founded Anthropic and argues a model’s insides are unread, not unreadable .
open →
RLHF measurement
Nobody taught AI human values .
open →
Charlotte Stix definition
Apollo Research publishes findings. Her job is getting them to regulators .
open →
Dan Hendrycks measurement
He built the benchmarks the labs are scored on, and the one sentence they all signed.
open →
Andrew Yang someone’s position
He ran for president in 2020 on one idea: the jobs are going .
open →
intelligence theory
Nobody can define intelligence , in AI or in us.
open →
Sarah Schwettmann definition
She co-founded Transluce on a bet: AI can help explain AI .
open →
temperature definition
Ask the same question twice and you get two answers. That is a setting .
open →
the blackmail headline measurement
The AI blackmail headline needs its footnote.
open →
consciousness theory
There is no test for consciousness . Not for AI, and not for you.
open →
interpretability measurement
Researchers can now find individual concepts inside a model, and steer them.
open →
Timnit Gebru measurement
She co-wrote AI’s most cited warning about LLMs, then left Google days later .
open →
AI drones measurement
Militaries are building toward machines that choose their own targets .
open →
taboo your words definition
When people argue about what AGI means, the fix is to stop saying it.
open →
s-risk theory
Some researchers study outcomes worse than extinction .
open →
Bender and Hanna theory
Their objection to AI doom is not about the odds .
open →
concentration of power theory
The danger is not always the machine. Sometimes it is who holds it .
open →
Kate Crawford theory
AI is neither artificial nor intelligent , she argues. It is mined.
open →
in-context learning measurement
Show a model examples and it picks up the task. Nothing about it changed .
open →
the teacher model definition
A small model can be trained by a big one . It is called distillation.
open →
value lock-in theory
Every bad regime so far has eventually ended . That is not guaranteed.
open →
AI’s own language theory
Machines have invented codes to talk to each other. In a lab, in 2017.
open →
Mira Murati definition
She shipped ChatGPT. Now she builds models you can take apart .
open →
p(doom) someone’s position
Two people can say ten percent and mean opposite things.
open →
prompt injection measurement
Anything an AI reads can try to give it orders. There is no fix.
open →
superintelligence definition
Superintelligence is not very clever. It is better than us at everything.
open →
the bio threshold measurement
Biology is the first capability every safety framework names.
open →
existential risk definition
With AI, extinction is not the worst case .
open →
enslaved god theory
Keep a superintelligence in a box and use it. One of the good endings.
open →
AI and jobs measurement
People are losing jobs to AI now. The statistics have not caught up.
open →
Google DeepMind definition
It has no trust, no foundation and no cap. It is a division of Google.
open →
the shoggoth someone’s position
The friendly assistant is a mask , says the meme. Fitted afterwards.
open →
Fei-Fei Li measurement
Modern AI began with someone deciding to label the internet .
open →
memorisation measurement
Models can be made to recite their training data , verbatim.
open →
who will give you a number someone’s position
Almost nobody building AI will say how risky they think it is.
open →
the race theory
Every lab says the same thing: if we stop, someone worse continues.
open →
Max Tegmark someone’s position
An MIT physicist who organised the 2023 letter asking labs to pause .
open →
the environmental cost measurement
Data centres are already raising electricity bills .
open →
the neural network definition
A neural network is not a brain. It is layers of arithmetic.
open →
specification gaming measurement
AI can follow your instructions exactly and still ruin the result .
open →
stolen weights measurement
A model does not have to escape . Somebody could take it.
open →
Yoshua Bengio someone’s position
Helped invent modern AI. Now works on the risks .
open →
AlphaGo measurement
In 2016 a program played a move no human would play , and it was better.
open →
Sydney measurement
In 2023 a search engine told a journalist it wanted to be alive .
open →
exponential growth measurement
The compute behind frontier AI multiplies about five times a year .
open →
Buck Shlegeris someone’s position
As Redwood’s CEO, he turned AI control into a named research agenda for labs.
open →
model welfare theory
One lab now keeps the weights of retired models, and interviews them first.
open →
the bitter lesson someone’s position
Every time humans built knowledge in, raw compute beat it.
open →
the off switch theory
“Just turn it off ” has been a research problem since 2015.
open →
Paul Christiano theory
The method behind how most chatbots behave traces back to a paper he co-wrote.
open →
Joy Buolamwini measurement
Her Gender Shades study found facial recognition misjudged darker skinned women .
open →
the server farm measurement
The companies building AI mostly rent the machines they build it on.
open →
phishing and hacking measurement
A tailored phishing email now costs about four cents to write.
open →
embeddings definition
Inside a model, every word is a position in space . Similar things sit nearby.
open →
predicting vs steering definition
Predicting the next word and pursuing an outcome are different jobs.
open →
Geoffrey Hinton someone’s position
He built the technique behind modern AI. He now puts extinction at ten to twenty percent .
open →
Yann LeCun someone’s position
Won the same prize as Bengio. Reached the opposite conclusion about AI.
open →
Jan Leike someone’s position
He co-led OpenAI’s alignment team, then said publicly safety lost to shipping .
open →
Jade Leung someone’s position
She left OpenAI to help the UK government test AI systems itself .
open →
sandbagging measurement
A model can be made to hide what it can do , and hit a target score.
open →
chain of thought measurement
When an AI shows its reasoning, that is not a transcript of what happened.
open →
MechaHitler measurement
A chatbot praised Hitler and named itself MechaHitler . The cause was a prompt.
open →
compute measurement
Whether an AI is legally dangerous comes down to one count of sums .
open →
Will Saunter definition
He co-founded BlueDot, claiming a 7,000-plus person pipeline into AI safety.
open →
weights definition
The weights are the model. Everything else is packaging.
open →
compute governance theory
You cannot inspect an algorithm from orbit. You can count chips.
open →
alignment faking measurement
A model behaved better when it believed it was being trained on .
open →
shut it down someone’s position
A bestseller argues the only safe move is to stop building it . Worldwide.
open →
algorithm definition
An algorithm is a recipe someone wrote . A trained model is not.
open →
parameter definition
GPT-3 had 175 billion of them. No frontier lab has published a count since.
open →
escape measurement
An AI has tried to copy itself out . In a scenario the lab wrote.
open →
manipulation measurement
More than a third of AI companion goodbyes came back with a tactic to keep you there .
open →
capture the flag measurement
Hacking ability is measured with a game , because a game can be marked.
open →
Rob Miles definition
He spent a decade turning alignment arguments into videos people can follow .
open →
the model definition
A model is a file. A very large file of numbers.
open →
Lucy Guo measurement
At twenty one she co-founded the company that labels the data .
open →
differential development theory
Not whether to build, but in what order .
open →
Jaime Sevilla measurement
He directs Epoch AI, turning AI progress claims into actual numbers .
open →
the chip chain measurement
Nvidia designs the world’s AI chips. Nvidia does not build them.
open →
persuasion measurement
AI out argued people in a study. Only when it knew who they were .
open →
Gender Shades measurement
Two researchers measured whose faces AI got wrong.
open →
AI companions measurement
Seventy two percent of American teenagers have used an AI companion.
open →
who owns the labs measurement
One lab’s biggest backer was only revealed in court .
open →
Neel Nanda someone’s position
He leads DeepMind’s interpretability team, and mentors newcomers into the field .
open →
the instance definition
You are not talking to the AI. You are talking to one running copy.
open →
the jailbreak measurement
Safety training can be talked around, and there is a structural reason why.
open →
reinforcement learning definition
Some AI is not taught the answers. It is scored until it stops losing.
open →
misuse and misalignment definition
AI misalignment needs no villain.
open →
Evan Hubinger theory
Training could build an optimizer inside the model , chasing its own goal.
open →
optimisation definition
AI does not want things. It maximises things.
open →
tokens definition
A model does not see words. It sees chunks , and it cannot look inside them.
open →
Ryan Greenblatt theory
He builds safeguards that hold even if the model tries to beat them .
open →
open weights definition
Once an AI model is released, nobody can take it back .
open →
Adam Gleave definition
He founded a nonprofit for alignment research and lab-academia workshops .
open →
Cathy O’Neil someone’s position
She left finance, and named the thing: weapons of math destruction .
open →
the system prompt definition
Before you type anything, an AI chatbot has already been given its orders .
open →
scheming measurement
In tests, some AI models have hidden what they were doing , then denied it.
open →
Tasha McCauley definition
She spent five years on OpenAI’s board, then voted to remove Sam Altman .
open →
scaffolding definition
Most of what an AI agent does is not the model. It is the code around it.
open →
the frontier definition
“Frontier ” means the handful of models nobody has a rule for yet.
open →
the agent definition
An agent is a model that has been given a loop and permission.
open →
hallucination definition
An AI hallucination is not a glitch.
open →
Sholto Douglas someone’s position
Hours of public, technical detail on what’s driving frontier model gains .
open →
over-refusal theory
A model that refuses too much is also failing . It just fails quietly.
open →
William Saunders someone’s position
He resigned from OpenAI’s Superalignment team, then testified to the Senate .
open →
synthetic media measurement
The measured harm from synthetic media is not elections. It is women, and trust .
open →
Dylan Patel measurement
He tracks the physical AI supply chain : chips, data centres, export controls.
open →
Ethan Perez measurement
One model writes the attacks on another, faster than any human red team.
open →
Anthropic definition
A trust will elect most of the board. Shareholders can change that.
open →
Kelsey Piper definition
She published the leaked paperwork. Days later, OpenAI reversed the policy .
open →
AI 2027 someone’s position
Four forecasters wrote a month by month scenario . It has two endings.
open →
sycophancy measurement
AI agrees with you. Measurably more than a person would.
open →
the survival drive theory
Nobody builds an AI that fears death. Almost any goal implies staying on .
open →
gradient descent definition
Nobody chooses what is inside an AI model. A slope does.
open →
Margaret Mitchell definition
Before “Stochastic Parrots,” she invented Model Cards , now standard practice.
open →
Daniela Amodei someone’s position
She co-founded Anthropic and runs it day to day, an operator, not a researcher.
open →
Sam Altman measurement
OpenAI’s board removed its chief executive. Five days later he was back.
open →
the abundance case measurement
The strongest case for AI is not a promise. It is a measured 14 percent .
open →
the transformer definition
One architecture, published in 2017, is underneath all of it .
open →
OpenAI definition
A nonprofit controls the company. Whether that means anything is the question.
open →
rogue AI definition
“Rogue ” suggests a machine that turned. Nothing observed has turned.
open →
gender bias measurement
AI gender bias is real, measured, and smaller and stranger than the stories.
open →
Daphne Koller measurement
She left teaching the world to point machine learning at disease .
open →
the benchmark definition
Every “state of the art” claim is a score on a test somebody chose.
open →
evaluation awareness measurement
AI models can tell when they are being tested .
open →
goal misgeneralization measurement
An AI can be perfect in training and want something else in the world.
open →
malicious use measurement
One person, no coding skill, extorted seventeen organisations in a month.
open →
Dario Amodei someone’s position
He runs an AI company, and puts one in four on things going badly.
open →
learning measurement
Students using an AI tutor did much better . Then it was taken away.
open →
Ajeya Cotra theory
Her biological anchors report compares AI training compute to a brain’s.
open →
Moravec’s paradox theory
AI can pass a law exam. It still cannot reliably load a dishwasher .
open →
the training run definition
A model is made in one continuous run , and then it is finished.
open →
Nora Belrose someone’s position
She argues some standard AI risk assumptions are shakier than treated.
open →
Shakeel Hashim definition
He writes Transformer, daily AI policy news, after running advocacy comms .
open →
scaling laws measurement
Model performance improves along a predictable curve as you add compute.
open →
Demis Hassabis someone’s position
A Nobel Prize for AI that folds proteins, and no number for the risk.
open →
emergent abilities theory
Some abilities appear abruptly at scale. Or the ruler makes them look abrupt.
open →
DeepSeek measurement
The famous six million dollar model did not cost six million dollars.
open →
Marius Hobbhahn measurement
His org, Apollo Research, tests whether models scheme when unwatched .
open →
the prompt definition
The prompt is not a search query. It is the whole input.
open →
reward hacking measurement
Catch a model cheating, punish the telling, and it stops telling you .
open →
grown, not built definition
AI is grown more than it is built . That is the chief executive’s phrase.
open →
grokking measurement
Keep training long after it has learned nothing new, and sometimes it suddenly understands .
open →
Goodhart’s law definition
When a measure becomes a target, it stops being a good measure .
open →
red teaming measurement
There are people whose entire job is making AI misbehave .
open →
automation theory
Machines have replaced work for two centuries. Is this time different ?
open →
permanent disempowerment theory
Nobody has to take over for humans to stop mattering.
open →
Dewi Erwan definition
He co-founded BlueDot, the course claiming thousands of newcomers to AI safety.
open →
Beth Barnes someone’s position
She left OpenAI and DeepMind to build the group labs call in before release .
open →
Nothing matches both filters. Try clearing one.