PopperWick: a research system where the models check each other

Python, Bash and four model families · Since July 2026 · Private repository, 1,064 commits as of September 2026

Why it exists

It started because these models kept annoying me. They are trained to hand you a complete answer every time, so each one is like a brilliant junior analyst who has never once said "I don't know." If there is a gap, he fills it in so smoothly you miss it. And their mistakes are not random; they all lean the same way, so asking a model to check its own work is worthless.

How it is arranged

One model runs the show. Cheaper models do the bulk reading. Models from other companies cross-examine the results, blind to each other when they are judging the same thing. When they disagree, that is the cheapest alarm I have. When they all agree, that is still not proof, because they might all have read the same wrong source, so the last step is always the original document, not another model's opinion.

How PopperWick is arranged The conductor sends work to the bulk readers, the cross-examiners and the verifier. Each of the three sends its verdict back to the conductor, and every verdict goes to the primary source, which has the last word. Conductor Bulk readers Cross-examiners Verifier Primary source verdicts the last word

What it runs on

My trading research, my classes and my job search. Every review gets saved with a verdict, so I can go back and see who was right. The best thing it has given me so far is a list of trading ideas I kept hearing about that did not survive testing. Not one did. I do not trade any of it; that is a standing rule.

What it produces

The part I can show without the repository is the backtest bench: a grader that writes down its formulas and reads a trade log for evidence of an edge rather than for profit. A worked example on QuantPad's public demo log: grade D, with the charts that show why.

Built and not built

Built: the conductor, the reading and review lanes, the archive of verdicts, and a page where the models sit in one room, get the same question, and I read them side by side. One session on that page, replayed: three models on a probability puzzle, and when they reviewed each other's answers without names, all three caught the same wrong number, including the model that wrote it. Not built: roles, where one argues for, one argues against, and one only hunts for what could go wrong. One day I want the whole thing on a machine in my apartment, on models I run myself.