Frontier Designchallenges for AI

We distinguish model-graded from programmatic evaluation. VLM means a vision-language model is actually used to inspect or judge an image. Programmatic scoring uses deterministic image processing, geometry, masks, or export checks. The two benchmarks below use programmatic scoring.

Benchmarks + gapsTwo experiments in creative AI: blur fields and recovering editable layers.
ProgrammaticBlur and color fieldsImage model + programmaticLayer retrieval
How we scoreGenerated results are measured against known reference pixels and layers.