Alignment beyond human evaluation.
Palaestra Research is a research lab working on AI alignment. Our focus is the intersection between scalable oversight and character training, as well as understanding new emergent multi-actor environments.
Research
-
SEP 2026
LessWrong ↗
In the HF and German wiki swarm incidents, agents used public web services such as counters and encoded URLs to signal activity and relay upcoming evaluation questions and answers. Our reproductions show that frontier models readily set up covert channels like these, raising concerns about contamination in open-web evaluations.
-
SEP 2026
LessWrong ↗
MonitoringBenchHonest adds 550 benign trajectories paired with 550 attack trajectories from MonitoringBench, matched by model and task. The dataset supports evaluation of coding-agent monitors, including false-positive rates and decision thresholds.
-
MAY 2026
Read → · Preprint ↗ · LessWrong ↗
On code-correctness and logic tasks, letting two stronger models debate raises a weaker judge's labeling accuracy well beyond consultancy baselines — by 14–16 macro-F1 points for the strongest debater pairs. The gains appear precisely when the critic beats the judge as a standalone classifier and the judge treats the critique as a claim to verify rather than defer to.
Macro-F1 lift of debate over consultancy, paired by question, across judge/debater pairs. -
FEB 2026
Read → · LessWrong ↗
An inference-time study of generative debates, in which a model defends an answer it produced itself against an independent critic. Established our verdict-accuracy methodology and mapped the regimes where debate transcripts help and fail to help weak judges on coding and reasoning tasks.
-
2026
HuggingFace ↗
BigCodeBench+ v0.1.0: Cleaned coding benchmark for verdict-accuracy research
A remediation of BigCodeBench with cleaned task specifications and tests. Fixing the benchmark reduced label noise by 20–25 percent.
-
2026
HuggingFace ↗
Uncontaminated Math Olympiad 2026 (UCMO): recent olympiad problems for contamination-free reasoning evaluation
Olympiad problems collected from 2026 competitions, postdating current model training cutoffs, for reasoning evaluation free of data contamination.
Current team

Ethan Elasky
AI researcher previously supported by Coefficient Giving, studying scalable oversight, debate, and AI control. Graduated from UC Berkeley as a University Medal candidate for top graduating senior. Former researcher at Academia Sinica. Ethan is fluent in Mandarin.

Frank N.
Frank studied pure mathematics at UC Berkeley. Frank's focus is understanding the character stability of increasingly agentic AI as well as new emergent multi-actor environments.

Can Kucukkurt
CS graduate from Purdue with a current focus on eval development. Can is currently working on monitoring benchmarks.

Laurence Tarquinio
Institutional investor with experience across healthcare, real estate, and technology, now working in private credit and venture. Laurence advises Palaestra on operations and works on research into the economic and real-world effects of advanced AI.