Frontier Designchallenges for AI
We distinguish model-graded from programmatic evaluation. VLM means a vision-language model is actually used to inspect or judge an image. Programmatic scoring uses deterministic image processing, geometry, masks, or export checks. The two benchmarks below use programmatic scoring.
Benchmarks + gapsTwo experiments in creative AI: blur fields and recovering editable layers.
ProgrammaticBlur and color fieldsImage model + programmaticLayer retrievalHow we scoreGenerated results are measured against known reference pixels and layers.
