Φ-MPO Fair Multi-Level Preference Optimization

Agentic Multi-Turn Reasoning:
A Fairness Approach

Thanh-Dat Truong1, Sankalp Pandey1, Hugh Churchill2, Jackson Cothren3, Marios Savvides4, Khoa Luu1
1CVIU Lab, University of Arkansas 2Department of Physics, University of Arkansas 3Department of Geosciences, University of Arkansas 4Carnegie Mellon University

Accepted to NeurIPS 2026

Overview of Φ-MPO, comparing its benchmark performance and support for long-horizon reasoning and data imbalance with prior approaches.
Learning to reason across multiple turns, with supervision for intermediate decisions and balanced emphasis on hard reasoning pairs.

Highlights

  • Multi-level preference learning: combine trajectory-level preferences with aligned state-level comparisons to improve credit assignment over long reasoning horizons.
  • Fairness under data imbalance: focal weighting reduces the influence of frequent, easy pairs and preserves learning from hard, informative reasoning decisions.
  • Theoretical analysis: characterize the shared pairwise optimization direction of DPO and GRPO under the paper’s assumptions, and show how the fair objective suppresses easy-majority dominance.
  • Strong results across four task families: Φ-MPO improves accuracy over Flow-GRPO on all ten reported benchmarks with the Qwen2.5-7B-Instruct backbone.

Abstract

Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or Φ-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.

The Φ-MPO Framework

The agent coordinates an Action Planner, Tool Executor, Execution Verifier, and Solution Generator through shared, evolving memory. Φ-MPO trains the planner from preferences over both complete trajectories and aligned intermediate reasoning states.

The Φ-MPO framework: multi-turn agentic interaction, preference construction, and optimization using trajectory-level DPO and fair multi-level preferences.
The proposed Φ-MPO framework. Select any figure to view it at full resolution.

From final outcomes to intermediate decisions

Trajectory-level DPO encourages successful complete rollouts. Multi-Level Preference Optimization adds comparisons between semantically aligned states in preferred and dispreferred trajectories, providing denser supervision for planning and tool use.

Balancing easy and hard preference pairs

The fair objective weights each aligned pair by (1 − p)γ, where p is its preference probability. This suppresses easy pairs that already have high preference probability. At γ = 0, the objective reduces to standard MPO; the experiments use γ = 2.

Why Data Imbalance Matters

Topics, difficulty levels, and reasoning patterns occur at uneven frequencies in agentic training data. Dominant patterns produce many redundant trajectories, while successful rollouts for rare or difficult cases remain scarce. Φ-MPO addresses this imbalance during preference optimization.

Analysis of the DeepMath data distribution and the influence of data imbalance on model performance.
The influence of imbalanced data, illustrated through the DeepMath analysis in the paper.

Experimental Results

Evaluation spans knowledge-intensive search, agentic reasoning, mathematical reasoning, and scientific reasoning. With Qwen2.5-7B-Instruct, Φ-MPO reaches 81.6% on Bamboogle, 41.7% on GAIA, 50.0% on AIME24, and 86.0% on MedQA.

Search and agentic reasoning · Accuracy (%) ↑
MethodBackbone / sizeBamboogle2WikiHotpotQAMusiqueGAIA
Qwen2.5-7B-Inst7B12.023.021.06.03.2
Qwen2.5-14B-Inst14B21.626.720.08.05.5
Qwen2.5-32B-Inst32B24.026.727.06.09.5
Llama-3.3-70B-Inst70B18.422.752.016.03.2
GPT-4o-mini—40.835.641.015.07.1
GPT-4o—68.849.554.024.017.3
SFT7B-Inst12.025.922.06.63.2
Iter-RetGen7B-Inst36.833.637.417.83.9
Search-R17B-Inst43.238.237.014.619.1
ZeroSearch7B-Base27.835.234.618.016.5
ReSearch7B-Base42.447.643.522.317.3
StepSearch7B-Base40.036.638.622.6-
VerlTool7B-Base46.445.344.819.311.2
AutoGen7B-Inst59.644.050.015.96.3
AgentFlow7B-Inst58.460.051.319.217.2
FlowGRPO7B-Inst69.677.257.025.333.1
Φ-MPO7B-Inst81.681.569.035.041.7
Φ-MPOQwen3.5-9B85.685.077.040.049.6
Mathematical and scientific reasoning · Accuracy (%) ↑
MethodBackbone / sizeAIME24AMC23GameOf24GPQAMedQA
Qwen2.5-7B-Inst7B6.747.533.034.066.0
Qwen2.5-14B-Inst14B6.760.025.031.075.0
Llama-3.3-70B-Inst70B6.747.531.035.067.0
Llama-3.1-405B-Inst405B26.747.523.030.062.0
GPT-4o-mini—13.357.516.027.066.0
GPT-4o—13.360.032.031.060.0
SFT7B-Inst6.747.533.034.066.0
SimpleRL7B-Base16.760.033.045.065.0
Open-Reasoner7B-Base16.754.932.034.054.0
General-Reasoner7B-Base13.355.033.035.561.0
Luffy7B-Inst30.744.833.034.077.0
TIR7B-Inst10.050.033.042.076.8
ToRL7B-Inst20.060.031.035.076.5
AutoGen7B-Inst13.357.524.042.072.0
AgentFlow7B-Inst16.747.431.037.076.0
Flow-GRPO7B-Inst40.061.553.047.080.0
Φ-MPO7B-Inst50.072.565.059.086.0
Φ-MPOQwen3.5-9B73.380.075.064.090.0

All accuracy values are reproduced from the paper. 7B-Inst: Qwen2.5-7B-Instruct; 7B-Base: Qwen2.5-7B-Base. A dash indicates an unreported score or an omitted proprietary model size. Tables scroll horizontally on smaller screens.

What Each Objective Contributes

Adding multi-level supervision improves trajectory-level DPO. Adding the fairness objective improves accuracy further on all four ablation benchmarks.

Objective ablation · Qwen2.5-7B-Instruct · Accuracy (%) ↑
Training objectiveBamboogle2WikiGAIAAIME24
Trajectory-level DPO61.661.521.330.0
+ Multi-level preferences (MPO)76.076.036.243.3
+ Fairness weighting (Φ-MPO)81.681.541.750.0

Multi-Turn Reasoning in Action

Case study from the paper showing a multi-turn reasoning example for Φ-MPO.
A qualitative case study from the paper.

BibTeX

@inproceedings{truong2026agentic,
  title={Agentic Multi-Turn Reasoning: A Fairness Approach},
  author={Truong, Thanh-Dat and Pandey, Sankalp and Churchill, Hugh and
          Cothren, Jackson and Savvides, Marios and Luu, Khoa},
  booktitle={Advances in Neural Information Processing Systems},
  year={2026}
}