DailyChatSunday · Sep 20

Put AI claims to the test.

3 stories · 2 minute read
01 / AI & SCIENCE

A biology lab gives AI a physical test

Anthropic confirmed a Bay Area biology lab to Reuters. The company says it is not running clinical trials; using Claude to direct robotic experiments remains early work, with human oversight essential.

Read the source ↗ english.aaj.tv
02 / AI RESEARCH / COMPANY METRICS

AI-led research still includes human supervision

Anthropic reports Claude led 26% of its AI R&D work in August under human supervision, with none fully autonomous. Its prototype index uses Claude to assess internal work; independent verification and comparable cross-lab methods remain unresolved.

Read the source ↗ www.anthropic.com
03 / RESEARCH PRACTICE / OPINION

Test predictions against evidence outside the model

Two independent checks: compare output with a known answer, or compare a prediction with a physical measurement.
Conceptual illustration of the essay’s distinction. Agreement in a test does not establish reliability in every setting.

Deep Manifold draws a lesson from numerical engineering: known-answer tests check computations, while physical experiments test predictions against independent observations. The essay argues that plausible outputs alone cannot establish reliability in science or engineering.

Read the source ↗ deepmanifold.substack.com