Research
Benchmarks and studies from the Fluxyr team.
Fluxyr Research · Code generation
Read the report →Quill: Web Apps as Micro-Prompts
A proof-of-concept for ephemeral code generation — one spec, three languages.
Technical benchmark · Coding models
Read the report →Evaluating top-tier coding models for ephemeral code generation
Ten frontier models, forty single-purpose tasks, executed and graded on whether the code actually runs, actually works, and what it costs. The finding: for small programs, correctness is nearly solved — and a cheap model plus one self-review pass matches the best models at a fraction of the price.