AI judges
Four AI specialists that review the agents from different angles. They vote before big decisions. You control what each judge cares about, and how loud its vote is.
Pick an experiment. Each enabled judge writes a short verdict. Their weighted vote becomes the overall panel decision.
Looks for hidden danger. Rejects strategies that rely on luck, thin data, or unrealistic assumptions.
Model: GPT-4o-mini. Prompt changes take effect on the next review round.
Challenges good results. Assumes agents might have gotten lucky — and looks for proof they didn't.
Model: GPT-4o-mini. Prompt changes take effect on the next review round.
Watches the health of the whole learning process. Suggests more variety when things get stale.
Model: GPT-4o-mini. Prompt changes take effect on the next review round.
Forms hypotheses about what makes strategies work. Suggests new experiments.
Model: GPT-4o-mini. Prompt changes take effect on the next review round.