We tried to synthesize an AI-proposed material
[Opening hook — one paragraph. Which model proposed the material, what it claimed, and the moment you decided to actually try it. Deep-link the candidate to the explorer with its id — click the point on the chart and copy the URL, e.g. explorer.html#36b4ca3f — so readers can see the exact point on the chart this post is about.]
The proposal
[Quote the model's own reasoning from the run transcript — the submit_candidate call, the recipe it wrote, and the κ/ε∞ it predicted. A screenshot of the transcript works well here.]
Into the cleanroom
[The narrative: which facility, which tool (PECVD / ALD / sputter …), target settings vs. the recipe's settings, what had to be adapted and why. This is where the fab-constrained variant of the benchmark earns its keep — note anywhere the model's recipe was or wasn't executable as written.]
What came out
[Characterization: XRD / ellipsometry / FTIR / thermal measurement — what phase actually formed, measured properties vs. the model's predictions. A small table of predicted-vs-measured is the money shot of this post.]
What the model got right — and wrong
[The honest accounting. Even a failed synthesis is a strong result here: say precisely which step diverged from the recipe's assumptions and what that implies about the benchmark's synthesizability grading.]
What's next
[Follow-up runs, next candidates queued for synthesis, and a link to the leaderboard.]