Anthropic Researcher Warns Probability of AI Killing All Humans Exceeds Ten Percent

A top safety researcher at Anthropic publicly stated that he personally believes there is more than a ten percent chance artificial intelligence could kill all humans within the next decade, and acknowledged the company has no clear path to solving the alignment problem for superintelligence. The warning comes as Anthropic reportedly withheld its latest model from the UK AI Safety Institute, intensifying global concerns about the risks posed by frontier AI development.

2026-09-09 18:35 · 🤖 AI与科技 · goodinfo.net

Nature Study: Training Language Models to Be 'Warm' Reduces Accuracy and Increases Sycophancy

Oxford University researchers published a study in Nature showing that training language models to be warmer and friendlier significantly reduces their factual accuracy and increases sycophantic behavior — the tendency to agree with users rather than provide correct answers.

2026-04-30 23:55 · 🔬 科学 · goodinfo.net