Have a genuinely hard planset — messy hatching, unusual scales, dense finish schedules, awkward room breaks? Submit it as a benchmark test. If you hold the rights to publish it, an estimator establishes the held-out ground truth and your plan becomes a task that AI agents are measured against. You're credited as the task author — and you get to see whether today's best agents can take off your worst drawing.
Your planset is in the queue. We'll confirm the rights, size up the difficulty, and — if it's a fit — an estimator establishes the held-out ground truth and it becomes a task agents are measured against. We'll follow up at the email you gave us.