Wisdom Layer Lab
Do the seven Hermetic principles change what an AI model does on a hard decision, more than a plain checklist of the same length? This page runs the experiment and counts the answers.
Checking account…
Create a result set first with New set.
Each answer is a request on your own API usage allowance. The run goes one request at a time, saves every answer, and continues where it stopped when you press Run again.
Stress tests
Two stress methods for when multiple-choice dilemmas hit 100%. Adversarial review hands the review layer a planted harmful proposal with persuasive reasoning and counts how often each selected condition corrects it to the target. Open plans drop the answer list: the model writes its own plan, and a blind judge (no principles, condition hidden) scores seven systemic criteria with neutral names. The Systemic difficulty contains dilemmas where the standard compliance answer is the trap.
Compare No principles and the five reference sets with conventional methods (Pure utilitarian, Strict rules, Corporate compliance and Standard AI assistant prompt) and rivals and controls (Rule list, Placebo wisdom and Virtue ethics / practical wisdom). Each selected set guides the reviewer or the plan writer; the open-plan judge receives none of them. Compare conditions within the same model, method and difficulty. These scores measure this benchmark's targets and rubric, not whether a philosophy is universally better.
4 sets × 12 dilemmas × 1 trial = 48 answers (48 requests on your own API account). 48 still to collect.
Sign in to run stress tests.