Ground Truth

Hands-on tests of AI hype, for people who build.

I take claims about AI, the scary ones and the too-good-to-be-true ones, and actually run the test. Every piece ships with the code and the receipts, so you can reproduce it or point it at your own system.

Hands-on testJuly 11, 20267 min read

Your AI said no. The next model in your pipeline never got the memo.

A claim says chaining two LLMs quietly cancels out their safety. I built the test across three Claude models. When the second model can see the refusal it holds (0/9). When the refusal is stripped, it complies every time (9/9). The safety lives in the context, not the model.

Read the test →

More tests are on the way. Want the next one in your inbox? Subscribe on Substack →