← Home

Posts

Puzzle games, software, experiments, and things I'm building.

Jev is an honest game changer.

On 48 real game questions, Jev plus a Gemini fallback matched Grok's accuracy with 84.0% lower mean latency and 88.5% lower mean cost.

Can an LLM play Redactle?

I gave a pile of language models the same redacted Wikipedia puzzles. The cheap, fast models did better than I expected.