Alignment beyond human evaluation.

Palaestra Research is a research lab working on AI alignment. Our focus is the intersection between scalable oversight and character training, as well as understanding new emergent multi-actor environments.

Research

  • SEP 2026

    Encoded Coordination on the Open Web

    Ethan Elasky, Can Kucukkurt, Frank N., David Africa

    LessWrong  ↗

    In the HF and German wiki swarm incidents, agents used public web services such as counters and encoded URLs to signal activity and relay upcoming evaluation questions and answers. Our reproductions show that frontier models readily set up covert channels like these, raising concerns about contamination in open-web evaluations.

  • SEP 2026 LessWrong  ↗

    MonitoringBenchHonest adds 550 benign trajectories paired with 550 attack trajectories from MonitoringBench, matched by model and task. The dataset supports evaluation of coding-agent monitors, including false-positive rates and decision thresholds.

  • MAY 2026

    Debate Helps Weak Judges Reward Stronger Models

    Ethan Elasky, Frank N.; Naman Goyal

    Read  →  ·  Preprint  ↗  ·  LessWrong  ↗

    On code-correctness and logic tasks, letting two stronger models debate raises a weaker judge's labeling accuracy well beyond consultancy baselines — by 14–16 macro-F1 points for the strongest debater pairs. The gains appear precisely when the critic beats the judge as a standalone classifier and the judge treats the critique as a claim to verify rather than defer to.

    Forest plot of macro-F1 lift of debate over consultancy across five model pairs on logic and code tasks
    Macro-F1 lift of debate over consultancy, paired by question, across judge/debater pairs.
  • FEB 2026 Read  →  ·  LessWrong  ↗

    An inference-time study of generative debates, in which a model defends an answer it produced itself against an independent critic. Established our verdict-accuracy methodology and mapped the regimes where debate transcripts help and fail to help weak judges on coding and reasoning tasks.

  • 2026

    BigCodeBench+ v0.1.0: Cleaned coding benchmark for verdict-accuracy research

    Palaestra Research

    HuggingFace  ↗

    A remediation of BigCodeBench with cleaned task specifications and tests. Fixing the benchmark reduced label noise by 20–25 percent.

  • 2026

    Uncontaminated Math Olympiad 2026 (UCMO): recent olympiad problems for contamination-free reasoning evaluation

    Palaestra Research

    HuggingFace  ↗

    Olympiad problems collected from 2026 competitions, postdating current model training cutoffs, for reasoning evaluation free of data contamination.

Current team