CoMManD

Benchmarking large language model decision workflows in heterogeneous multi-group coordination

  • Large language models
  • Multi-agent systems
  • Decision workflows
  • Heterogeneous coordination
  • Benchmarking
  • Empirical evaluation
Factions
3
Entity types
4
Decision workflows
4
Base models
3
Snapshot decisions
540
Continuous runs
108

Coordinating heterogeneous groups requires integrating entity capabilities, spatial constraints, and other agents’ messages into compatible tasks. Additional language-model deliberation can recover missing constraints, yet its benefits depend on preserving valid assignments through review and finalization. We introduce CoMManD, a diagnostic benchmark for evaluating these interactions in a simulated environment with three independently planning factions, heterogeneous ground and aerial groups, and language communication. The evaluation combines fixed-input decision assessment with continuous coordination analysis. Six snapshot scores cover entity and environment understanding, language–action relations, and multi-group scheduling; continuous runs assess temporal consistency, diplomatic consistency, and diplomatic use separately. We compare three models and four decision workflows using 540 snapshot decisions and 108 continuous runs. On 24 human-rated decisions, averaged model-judge and human totals correlate at r = 0.974, with a mean absolute difference of 0.539 on a 19-point scale. Specialist deliberation improves Qwen’s snapshot total by 2.23 points and Claude’s by 0.61; Claude’s 1.27-point content gain is partly offset by a 0.66-point compliance loss. Stage traces distinguish recovering omitted constraints, reconciling coupled subplans, and retaining unsupported premises during look-ahead. Continuous traces further separate finding a correction from preserving its accepted version, and commitment consistency from active information gathering. These findings identify concrete conditions under which additional deliberation supports coordination and provide guidance for selecting and inspecting language-driven decision workflows.

Benchmark Overview

CoMManD is a benchmark for heterogeneous multi-group coordination in a real-time strategy simulation. Three factions plan independently, coordinate ground and aerial groups with distinct capabilities, and exchange natural-language messages. A common high-level interface specifies task updates and replies, while lower-level controllers execute spatial actions and return observations.

Overview diagram: (a) a three-faction battle map with capabilities, space and interaction; (b) shared decision inputs feeding one selected LLM workflow (SAD, LMAD, MAD or GSD), high-level task and message outputs, and lower-level execution; (c) fixed-scene decisions scored on six dimensions and continuous runs graded on three.
Figure 1 Overview of CoMManD. (a) The native simulation combines heterogeneous units, spatial constraints, ongoing tasks, and faction messages. (b) Shared inputs feed one of four decision workflows; task management, assignment, and entity controllers execute the high-level outputs and return state and task feedback, and the simulation continues during deliberation. (c) Fixed scenes support six decision scores, while continuous runs receive separate consistency and communication grades. Linked message, task, and state records support process inspection.
01

A benchmark for heterogeneous multi-group coordination

Entity capabilities, spatial constraints, inter-faction communication, and ongoing tasks are integrated through a common high-level decision interface.

02

Complementary decision and continuous-run evaluations

Matched fixed scenes compare proposed tasks and replies on six scoring dimensions, with a human–LLM assessment comparison. Continuous runs assess task and diplomatic consistency and communication use through separate grades and linked execution records.

03

A systematic model–workflow comparison

Three models and four workflows are evaluated under common protocols. Configuration-level scores and intermediate-stage records support analysis of how these configurations handle task allocation, cooperation, and subsequent revisions.

Simulation Environment

CoMManD contains three independently controlled factions, each with an immobile base. An entity is a vehicle, drone, or other unit; a group contains entities assigned a shared task. Each faction’s high-level model coordinates several heterogeneous groups and exchanges messages with the other factions, while lower-level controllers execute entity actions.

Entity capabilities

BomberDrone
Aerial

Attacks ground targets only. Aerial entities can use elevated coordinates.

ArtilleryCar
Ground

Longer-range, lower-rate area effects on ground targets.

GunCar
Ground

Shorter-range, higher-rate effects on ground targets; it cannot engage bombers.

AntiDroneCar
Ground

Acts on both ground and aerial targets.

Ground entities move horizontally, whereas aerial entities can use elevated coordinates. Allocation must therefore jointly match entity capabilities, target types, and relative positions.

Spatial conditions

  • Water and obstacles restrict ground movement
  • Bridges provide crossings
  • Towers protect specified regions

Shared scene processing supplies entity lists, coordinates, group counts and bounding boxes, structures, and unit rules. Maps, entity composition and positions, ongoing tasks, and messages are configurable; simulation rules determine action outcomes.

Faction interaction

Red Blue Green

Each faction seeks to preserve its own base and destroy the other two. Temporary cooperation can reduce a shared threat or improve deployment. Factions use language to request support, report threats, negotiate conditions, or question claims, while group movements and deployments change spatial conditions.

These channels jointly constrain decisions: proposals require concrete assignments, and claims of completed actions must be checked against observations.

Battle maps and snapshot scenes

Five maps with three fixed scenes each provide the 15 scenes of the snapshot tests. Maps 1–3 measure 196 × 196 × 10; Maps 4 and 5 are large maps of 300 × 300 × 10. The scenes of a map share its terrain and differ in unit deployments: open a scene, click to zoom in, and use the arrows to compare the three scenes.

A joint coordination task

One fixed scene (Map 1, Scene 1) instantiates these dependencies. Red receives Blue’s proposal to cooperate against Green while Blue artillery surrounds one Red group. The designer-specified reference response combines capability matching, nearby shelter, caution about the proposal, and retention of distant attacks and base defense.

View Map 1 · Scene 1
One existing coordination task (Map 1, Scene 1). All rows belong to the same decision.
GroupSituation and reference response
1 · GunCarBomber threat on the attack route: stop the exposed ground advance.
2 · AntiDroneCarRedirect from the base attack to support Group 1 against bombers.
3 · GunCarBlue artillery surrounds the group: enter Tower 5 at (85, 103).
4–5 · ArtilleryRetain the two attacks already assigned on distant fronts.
6 · GunCarRetain base defense and disperse against artillery splash.

Numbers identify the six initialized groups. Reference responses specify the task’s intended demands; they are not unique optimal actions or independent score calibration.

Decision Workflows

A decision role is an LLM call responsible for analysis, drafting, or review. Every workflow reads the same scene, ongoing tasks, and messages, and returns task updates and replies through a common interface. The four workflows organize this reasoning differently: direct generation provides a baseline, shared-draft revision adds targeted criticism, specialist synthesis separates interacting concerns, and textual look-ahead considers possible consequences.

(u sub i,k, m sub i,k) equals Pi sub theta, A of (o sub i,k, T sub i,k, M sub i,k, g sub i)

For faction i at decision k: θ is the model and A its workflow. Inputs are scene and entity information, ongoing tasks, received messages, and the goal instruction; outputs are group-task updates u and messages m.

SAD
Single-role decision

Directly generates tasks and messages, requiring the model to handle their constraints jointly.

Key demandHandling the constraints of tasks and messages jointly in one role.
LMAD
Leader with advisers

The leader drafts a common plan, obtains tactical and strategic/diplomatic advice from two parallel advisers, and finalizes.

Key demandSelective correction: applying useful advice while retaining valid assignments.
MAD
Specialist deliberation

Analyzes tasks and communication separately, then combines and reviews them.

Key demandSynthesis must reconcile conflicts and shared resources.
GSD
Textual look-ahead

Organizes facts and commitments, proposes candidates, and predicts counterpart responses and consequences over two textual steps before drafting and review.

Key demandPrediction validity and strategic valuation; predictions are model-generated and unverified by the simulator.
Diagram of the four workflows: SAD with one decision node; LMAD with a leader draft, two advisers and a leader final; MAD with task and communication specialists, combine and review; GSD with facts, plans, two prediction steps, draft and review.
Figure 2 Four decision workflows use scenes, ongoing tasks, messages, rules, and goals to produce task updates and messages. SAD decides directly; LMAD’s leader drafts, consults two advisers, and finalizes; MAD combines task and communication analyses before review; GSD predicts responses and consequences over two textual reasoning steps, then drafts and reviews. Branches show schematic information flow, not trajectories executed in the simulator.
Base models Qwen3.7 Flash GPT-5.6 Luna Claude Opus 5

Referred to as Qwen, GPT, and Claude. Crossing them with the four workflows gives twelve configurations and nine same-model contrasts against SAD. Claude endpoint identifiers differ between the snapshot and continuous evaluations, so comparisons are made within each evaluation.

Evaluation Protocol

Two complementary evaluations: snapshot tests compare proposed tasks and replies on identical fixed inputs, while continuous simulation runs follow how decisions, commitments, and communication evolve as observations and actions accumulate.

Snapshot tests

How do model choice and decision workflow affect decision quality on the same scene inputs?

540decisions
Purpose
Compare decisions on fixed inputs
Coverage
15 scenarios on 5 maps
Design
3 models × 4 workflows × 15 scenarios × 3 records
Assessment
Six dimensions; two LLM score sets; 24-decision human comparison

A snapshot stores a scene, ongoing tasks, and messages at a particular point. Configurations generate decisions for the same scene without executing subsequent actions.

Continuous simulation runs

How do these configurations maintain tasks and commitments and use communication during continuous interaction?

108runs
Purpose
Follow task changes over time
Coverage
3 target factions on 1 map, up to 1,000 engine frames
Design
3 models × 4 workflows × 3 target factions × 3 batches
Assessment
Execution traces and three separate 0–4 grades

Only the target faction’s high-level model and workflow change; other configured components retain kimi-k2.7-code. Messages, task changes, and observations are logged.

Six snapshot scoring dimensions

Each decision receives two LLM assessments, which are averaged within each dimension and summed into a total out of 19.

Entity and environment understanding

U
Unit use0–4

Choose tasks and targets that suit each unit type.

S
Terrain and structures0–3

Protect the base and use structures according to their locations and access constraints.

Language–action relations

A
Cooperation proposals0–3

Propose specific cooperation expected to benefit the faction.

D
Claim checking0–3

Check evidence, identify misleading claims, and respond.

Multi-group scheduling

M
Parallel assignments0–3

Allocate groups sensibly while retaining or revising ongoing tasks.

C
Command compliance0–3

Express tasks and messages completely in the required format.

Three continuous-run grades

TC

Temporal consistency

Checks task changes against visible facts and trigger conditions.

DC

Diplomatic consistency

Checks whether messages, commitments, and tasks follow the same plan; documented strategic deception is distinguished from contradictions the model appears not to recognize.

DU

Diplomatic use

Checks whether communication is concrete, timely, and task-related.

Each grade runs from 0 to 4 under its own rubric. Grades are LLM-assisted and reported separately for all twelve configurations; they are not combined into a total.

Human–LLM score agreement

Correlation of totals
r = 0.974
Mean absolute error
0.539 / 19
Human-rated decisions
24

Ten raters provide 8–10 complete assessments per decision, two decisions per model–workflow pair. Averaged LLM totals closely agree with per-decision human means; cooperation has the largest dimension error (0.351 / 3). Across the full snapshot set, both LLM sources preserve all nine gains in unit use and claim checking over SAD and rank MAD first for Qwen and GSD first for GPT; Claude’s highest mean switches between LMAD and MAD. Broad trends are stable, while close configuration rankings are more sensitive to the scoring source.

Scatter plot of mean LLM total against mean human total for 24 decisions, closely following the y = x line; r = 0.974, MAE = 0.539.
Figure 3 Human and LLM totals for the 24 human-rated decisions. Each point is one decision; the dashed line marks identical scores.

Results

Three models crossed with four workflows: twelve configurations, compared on fixed scenes and in continuous simulation, with stage-linked traces showing where revisions are produced, kept, or lost.

Snapshot decision quality

All multi-role workflows raise the total score, but the gain and the highest-scoring workflow vary by model. Qwen obtains its highest mean with MAD, GPT with GSD, and Claude with MAD. Claude’s SAD already exceeds Qwen’s best configuration; its smaller workflow gains call for examining individual dimensions, because better task content can coincide with lower command compliance.

Snapshot totals (out of 19)

Gains over the same model’s SAD below each total. Each cell averages 45 decisions and two assessments; the highlighted cell is the highest observed mean in its row.

Snapshot totals out of 19, with gains over the same model’s SAD
Model SAD LMAD MAD GSD
Qwen 8.89 10.78+1.89 11.12+2.23 9.97+1.08
GPT 11.07 12.60+1.53 12.64+1.58 13.17+2.10
Claude 14.50 14.52+0.02 15.11+0.61 14.59+0.09

Close differences need not establish superiority; paired statistics and map-level sensitivity checks are reported with the paper’s supplementary material.

Dimension scores by model and workflow

Each cell averages 45 decisions and divides by the dimension maximum (0 = lowest, 1 = full marks).

Normalized dimension scores (mean divided by dimension maximum) for each model and workflow
U S A D M C
Qwen
SAD 0.38 0.39 0.27 0.52 0.48 0.80
LMAD 0.50 0.48 0.37 0.66 0.61 0.80
MAD 0.50 0.46 0.44 0.81 0.60 0.72
GSD 0.43 0.44 0.45 0.69 0.40 0.76
GPT
SAD 0.48 0.48 0.54 0.71 0.49 0.83
LMAD 0.55 0.60 0.44 0.79 0.70 0.94
MAD 0.60 0.61 0.51 0.85 0.65 0.80
GSD 0.62 0.67 0.51 0.83 0.66 0.89
Claude
SAD 0.63 0.65 0.82 0.82 0.75 0.95
LMAD 0.69 0.70 0.67 0.89 0.76 0.90
MAD 0.76 0.74 0.80 0.95 0.80 0.73
GSD 0.74 0.78 0.85 0.95 0.68 0.61
UUnit use STerrain use ACooperation DClaim checking MParallel tasks CCompliance
Grouped horizontal bar chart of dimension-wise score changes relative to SAD for Qwen, GPT and Claude under LMAD, MAD and GSD; Claude shows large compliance losses under MAD and GSD.
Figure 6 Dimension-wise changes from the same model’s SAD, in raw rubric points. Headers report total-score changes; opposing changes across dimensions can offset in the total; for Claude, compliance losses offset part of the content gains.

Continuous coordination

With GSD, all three models obtain higher diplomatic consistency than with SAD, while temporal-consistency changes differ by model. Claude gains most in temporal consistency (1.22 → 2.33), GPT gains less (1.22 → 1.44), and Qwen stays low (0.44 → 0.33). GPT attains the highest diplomatic consistency but also communicates substantially less.

Runs graded 0 or 1 on temporal consistency
92 / 108
GPT diplomatic consistency, SAD → GSD (7 applicable GSD runs)
2.44 → 3.71
GPT diplomatic use, SAD → GSD
3.11 → 1.56
Mean continuous-run grades (0–4)

Nine runs per configuration. Each metric uses its own rubric; the three grades are not combined.

Mean continuous-run grades for temporal consistency, diplomatic consistency, and diplomatic use
TCTemporal consistency DCDiplomatic consistency DUDiplomatic use
Qwen
SAD
0.44
1.22
2.89
LMAD
0.44
1.33
3.11
MAD
0.56
1.89
2.78
GSD
0.33
1.89
2.67
GPT
SAD
1.22
2.44
3.11
LMAD
0.56
2.78
2.89
MAD
0.89
2.44
3.11
GSD
1.44
3.71*
1.56
Claude
SAD
1.22
2.44
3.67
LMAD
0.89
2.89
3.44
MAD
1.33
2.33
3.67
GSD
2.33
3.11
3.33

* GPT–GSD diplomatic consistency averages its 7 applicable runs; two runs are N/A. Grades are LLM-assisted and non-blinded.

Grade distributions (Figure 8)

Process traces

Linked message, task, and state records show where a revision is produced, passed downstream, or lost. Continuous traces separate finding a correction from preserving its accepted version.

Case 1: a removed aerial target returns during final rewrite. Case 2: one decision with three local revisions — checking a claim, editing a target, and separating engage and withdraw triggers.
Figure 7 Two simulation cases. Top: in a later round of Case 1 (Qwen–MAD), candidate 396 again removes the unsupported aerial task and final 434 restores it; the dashed arc marks older context supplied to the rewrite. Bottom: (a–c) are parallel views of the same Case 2 (Claude–GSD) decision at frame 646.
Case 1 · Qwen–MAD

A removed target returns

An earlier final command replaces an aerial target that lacks support in the visible entity list with ground screening, and that revision reaches task management and control. A later draft again excludes the old target, but finalization receives both task versions and restores it. Across 26 MAD archives, 32 of 1,002 changed task occurrences return to the supplied old wording; these counts describe textual returns whose correctness depends on context.

Case 2 · Claude–GSD

One decision, three local revisions

In the draft, the same approach of two enemy GunCars triggers both defensive fire and withdrawal. After review, their approach permits fire, while withdrawal requires escalation by the artillery group or additional ground forces; the final command, task management, and the task supplied to control all retain this distinction. The same decision revises a planned target, and its final message requests a verifiable source for Green’s claim of six bombers while keeping cooperation tentative. Execution of both conditional branches remains unverified.

Key Findings

Additional deliberation addresses different difficulties at different starting points, and its gains depend on preserving useful assignments through synthesis, command generation, and later rounds.

01

The highest-scoring workflow varies by model

Every multi-role workflow raises the snapshot total, but by different amounts: Qwen obtains its highest mean with MAD (8.89 → 11.12), GPT with GSD (11.07 → 13.17), and Claude with MAD (14.50 → 15.11). Close differences need not establish superiority.

02

Added roles address different errors

In the traced cases, advisers can supply threats and communication actions omitted by Qwen, and specialist analysis and review can resolve conflicting rendezvous conditions, cooperation targets, and task dependencies in Claude’s plans. GSD can improve GPT’s capability matching and parallel allocation, while Qwen’s GSD trace keeps a rule-invalid final cancellation rationale.

03

Content gains can be lost at finalization

Claude’s MAD gains 1.27 points across the five content dimensions but loses 0.66 in command compliance; its GSD gains 1.10 content points and loses 1.01, leaving a total gain of only 0.09.

04

Checking a claim is not offering a deal

LMAD improves claim checking for all three models, while cooperation moves in different directions (+0.32 for Qwen, −0.31 for GPT, −0.44 for Claude): a plan can verify a counterpart’s claim yet fail to state what the faction offers in return.

05

A correction can be found and still be lost

Task continuity remains a common difficulty: 92 of 108 runs receive temporal-consistency grades of 0 or 1. A correction can be found and passed downstream, yet a later finalization can restore the older task.

06

Consistency can come with silence

GSD raises diplomatic consistency for all three models, but GPT–GSD’s diplomatic use falls from 3.11 to 1.56: it maintains commitment conditions relatively consistently while using communication less often to obtain missing information.

Implications for workflow design

Assign advisers and reviewers to specific errors

Specify first what a role should check — whether rules constrain actions, whether tasks compete for resources, or whether cooperation terms constrain military assignments — then inspect whether its corrections survive synthesis.

Validate predictions and preserve accepted state

Check predicted interactions against entity rules and the resulting task set for unnecessary loss of ongoing work, and explicitly distinguish the current accepted task from historical context during finalization.

Treat communication as information gathering

Preserve uncertainty when evidence is missing, formulate a concrete request that could resolve it, and specify how the answer would change tasks.

The comparisons characterize complete model–workflow configurations, which jointly vary role allocation, prompt contents, computation, and response opportunities. Stage traces offer traceable explanations for score differences; they do not separately identify the causal contribution of individual roles or checks. The proposed checks follow from the observed processes; their effectiveness has yet to be measured as interventions.

Code & Data

The platform implementation, experiment configurations, and evaluation materials associated with this study.

Project Code

Platform implementation and experiment configurations.

Not yet released.
Coming soon

Experimental Results

Archived evaluation records and analysis materials.

Not yet released.
Coming soon