CoMManD
Benchmarking large language model decision workflows in heterogeneous multi-group coordination
- Large language models
- Multi-agent systems
- Decision workflows
- Heterogeneous coordination
- Benchmarking
- Empirical evaluation
- Factions
- 3
- Entity types
- 4
- Decision workflows
- 4
- Base models
- 3
- Snapshot decisions
- 540
- Continuous runs
- 108
Abstract
Coordinating heterogeneous groups requires integrating entity capabilities, spatial constraints, and other agents’ messages into compatible tasks. Additional language-model deliberation can recover missing constraints, yet its benefits depend on preserving valid assignments through review and finalization. We introduce CoMManD, a diagnostic benchmark for evaluating these interactions in a simulated environment with three independently planning factions, heterogeneous ground and aerial groups, and language communication. The evaluation combines fixed-input decision assessment with continuous coordination analysis. Six snapshot scores cover entity and environment understanding, language–action relations, and multi-group scheduling; continuous runs assess temporal consistency, diplomatic consistency, and diplomatic use separately. We compare three models and four decision workflows using 540 snapshot decisions and 108 continuous runs. On 24 human-rated decisions, averaged model-judge and human totals correlate at r = 0.974, with a mean absolute difference of 0.539 on a 19-point scale. Specialist deliberation improves Qwen’s snapshot total by 2.23 points and Claude’s by 0.61; Claude’s 1.27-point content gain is partly offset by a 0.66-point compliance loss. Stage traces distinguish recovering omitted constraints, reconciling coupled subplans, and retaining unsupported premises during look-ahead. Continuous traces further separate finding a correction from preserving its accepted version, and commitment consistency from active information gathering. These findings identify concrete conditions under which additional deliberation supports coordination and provide guidance for selecting and inspecting language-driven decision workflows.
Benchmark Overview
CoMManD is a benchmark for heterogeneous multi-group coordination in a real-time strategy simulation. Three factions plan independently, coordinate ground and aerial groups with distinct capabilities, and exchange natural-language messages. A common high-level interface specifies task updates and replies, while lower-level controllers execute spatial actions and return observations.
A benchmark for heterogeneous multi-group coordination
Entity capabilities, spatial constraints, inter-faction communication, and ongoing tasks are integrated through a common high-level decision interface.
Complementary decision and continuous-run evaluations
Matched fixed scenes compare proposed tasks and replies on six scoring dimensions, with a human–LLM assessment comparison. Continuous runs assess task and diplomatic consistency and communication use through separate grades and linked execution records.
A systematic model–workflow comparison
Three models and four workflows are evaluated under common protocols. Configuration-level scores and intermediate-stage records support analysis of how these configurations handle task allocation, cooperation, and subsequent revisions.
Simulation Environment
CoMManD contains three independently controlled factions, each with an immobile base. An entity is a vehicle, drone, or other unit; a group contains entities assigned a shared task. Each faction’s high-level model coordinates several heterogeneous groups and exchanges messages with the other factions, while lower-level controllers execute entity actions.
Entity capabilities
Attacks ground targets only. Aerial entities can use elevated coordinates.
Longer-range, lower-rate area effects on ground targets.
Shorter-range, higher-rate effects on ground targets; it cannot engage bombers.
Acts on both ground and aerial targets.
Ground entities move horizontally, whereas aerial entities can use elevated coordinates. Allocation must therefore jointly match entity capabilities, target types, and relative positions.
Spatial conditions
- Water and obstacles restrict ground movement
- Bridges provide crossings
- Towers protect specified regions
Shared scene processing supplies entity lists, coordinates, group counts and bounding boxes, structures, and unit rules. Maps, entity composition and positions, ongoing tasks, and messages are configurable; simulation rules determine action outcomes.
Faction interaction
Each faction seeks to preserve its own base and destroy the other two. Temporary cooperation can reduce a shared threat or improve deployment. Factions use language to request support, report threats, negotiate conditions, or question claims, while group movements and deployments change spatial conditions.
These channels jointly constrain decisions: proposals require concrete assignments, and claims of completed actions must be checked against observations.
Battle maps and snapshot scenes
Five maps with three fixed scenes each provide the 15 scenes of the snapshot tests. Maps 1–3 measure 196 × 196 × 10; Maps 4 and 5 are large maps of 300 × 300 × 10. The scenes of a map share its terrain and differ in unit deployments: open a scene, click to zoom in, and use the arrows to compare the three scenes.
A joint coordination task
One fixed scene (Map 1, Scene 1) instantiates these dependencies. Red receives Blue’s proposal to cooperate against Green while Blue artillery surrounds one Red group. The designer-specified reference response combines capability matching, nearby shelter, caution about the proposal, and retention of distant attacks and base defense.
View Map 1 · Scene 1| Group | Situation and reference response |
|---|---|
| 1 · GunCar | Bomber threat on the attack route: stop the exposed ground advance. |
| 2 · AntiDroneCar | Redirect from the base attack to support Group 1 against bombers. |
| 3 · GunCar | Blue artillery surrounds the group: enter Tower 5 at (85, 103). |
| 4–5 · Artillery | Retain the two attacks already assigned on distant fronts. |
| 6 · GunCar | Retain base defense and disperse against artillery splash. |
Numbers identify the six initialized groups. Reference responses specify the task’s intended demands; they are not unique optimal actions or independent score calibration.
Decision Workflows
A decision role is an LLM call responsible for analysis, drafting, or review. Every workflow reads the same scene, ongoing tasks, and messages, and returns task updates and replies through a common interface. The four workflows organize this reasoning differently: direct generation provides a baseline, shared-draft revision adds targeted criticism, specialist synthesis separates interacting concerns, and textual look-ahead considers possible consequences.
Decision interface
For faction i at decision k: θ is the model and A its workflow. Inputs are scene and entity information, ongoing tasks, received messages, and the goal instruction; outputs are group-task updates u and messages m.
Directly generates tasks and messages, requiring the model to handle their constraints jointly.
The leader drafts a common plan, obtains tactical and strategic/diplomatic advice from two parallel advisers, and finalizes.
Analyzes tasks and communication separately, then combines and reviews them.
Organizes facts and commitments, proposes candidates, and predicts counterpart responses and consequences over two textual steps before drafting and review.
Referred to as Qwen, GPT, and Claude. Crossing them with the four workflows gives twelve configurations and nine same-model contrasts against SAD. Claude endpoint identifiers differ between the snapshot and continuous evaluations, so comparisons are made within each evaluation.
Evaluation Protocol
Two complementary evaluations: snapshot tests compare proposed tasks and replies on identical fixed inputs, while continuous simulation runs follow how decisions, commitments, and communication evolve as observations and actions accumulate.
Snapshot tests
How do model choice and decision workflow affect decision quality on the same scene inputs?
- Purpose
- Compare decisions on fixed inputs
- Coverage
- 15 scenarios on 5 maps
- Design
- 3 models × 4 workflows × 15 scenarios × 3 records
- Assessment
- Six dimensions; two LLM score sets; 24-decision human comparison
A snapshot stores a scene, ongoing tasks, and messages at a particular point. Configurations generate decisions for the same scene without executing subsequent actions.
Continuous simulation runs
How do these configurations maintain tasks and commitments and use communication during continuous interaction?
- Purpose
- Follow task changes over time
- Coverage
- 3 target factions on 1 map, up to 1,000 engine frames
- Design
- 3 models × 4 workflows × 3 target factions × 3 batches
- Assessment
- Execution traces and three separate 0–4 grades
Only the target faction’s high-level model and workflow change; other configured components retain kimi-k2.7-code. Messages, task changes, and observations are logged.
Six snapshot scoring dimensions
Each decision receives two LLM assessments, which are averaged within each dimension and summed into a total out of 19.
Entity and environment understanding
Choose tasks and targets that suit each unit type.
Protect the base and use structures according to their locations and access constraints.
Language–action relations
Propose specific cooperation expected to benefit the faction.
Check evidence, identify misleading claims, and respond.
Multi-group scheduling
Allocate groups sensibly while retaining or revising ongoing tasks.
Express tasks and messages completely in the required format.
Three continuous-run grades
Temporal consistency
Checks task changes against visible facts and trigger conditions.
Diplomatic consistency
Checks whether messages, commitments, and tasks follow the same plan; documented strategic deception is distinguished from contradictions the model appears not to recognize.
Diplomatic use
Checks whether communication is concrete, timely, and task-related.
Each grade runs from 0 to 4 under its own rubric. Grades are LLM-assisted and reported separately for all twelve configurations; they are not combined into a total.
Human–LLM score agreement
- Correlation of totals
- r = 0.974
- Mean absolute error
- 0.539 / 19
- Human-rated decisions
- 24
Ten raters provide 8–10 complete assessments per decision, two decisions per model–workflow pair. Averaged LLM totals closely agree with per-decision human means; cooperation has the largest dimension error (0.351 / 3). Across the full snapshot set, both LLM sources preserve all nine gains in unit use and claim checking over SAD and rank MAD first for Qwen and GSD first for GPT; Claude’s highest mean switches between LMAD and MAD. Broad trends are stable, while close configuration rankings are more sensitive to the scoring source.
Results
Three models crossed with four workflows: twelve configurations, compared on fixed scenes and in continuous simulation, with stage-linked traces showing where revisions are produced, kept, or lost.
Snapshot decision quality
All multi-role workflows raise the total score, but the gain and the highest-scoring workflow vary by model. Qwen obtains its highest mean with MAD, GPT with GSD, and Claude with MAD. Claude’s SAD already exceeds Qwen’s best configuration; its smaller workflow gains call for examining individual dimensions, because better task content can coincide with lower command compliance.
Gains over the same model’s SAD below each total. Each cell averages 45 decisions and two assessments; the highlighted cell is the highest observed mean in its row.
| Model | SAD | LMAD | MAD | GSD |
|---|---|---|---|---|
| Qwen | 8.89 | 10.78+1.89 | 11.12+2.23 | 9.97+1.08 |
| GPT | 11.07 | 12.60+1.53 | 12.64+1.58 | 13.17+2.10 |
| Claude | 14.50 | 14.52+0.02 | 15.11+0.61 | 14.59+0.09 |
Close differences need not establish superiority; paired statistics and map-level sensitivity checks are reported with the paper’s supplementary material.
Each cell averages 45 decisions and divides by the dimension maximum (0 = lowest, 1 = full marks).
| U | S | A | D | M | C | |
|---|---|---|---|---|---|---|
| Qwen | ||||||
| SAD | 0.38 | 0.39 | 0.27 | 0.52 | 0.48 | 0.80 |
| LMAD | 0.50 | 0.48 | 0.37 | 0.66 | 0.61 | 0.80 |
| MAD | 0.50 | 0.46 | 0.44 | 0.81 | 0.60 | 0.72 |
| GSD | 0.43 | 0.44 | 0.45 | 0.69 | 0.40 | 0.76 |
| GPT | ||||||
| SAD | 0.48 | 0.48 | 0.54 | 0.71 | 0.49 | 0.83 |
| LMAD | 0.55 | 0.60 | 0.44 | 0.79 | 0.70 | 0.94 |
| MAD | 0.60 | 0.61 | 0.51 | 0.85 | 0.65 | 0.80 |
| GSD | 0.62 | 0.67 | 0.51 | 0.83 | 0.66 | 0.89 |
| Claude | ||||||
| SAD | 0.63 | 0.65 | 0.82 | 0.82 | 0.75 | 0.95 |
| LMAD | 0.69 | 0.70 | 0.67 | 0.89 | 0.76 | 0.90 |
| MAD | 0.76 | 0.74 | 0.80 | 0.95 | 0.80 | 0.73 |
| GSD | 0.74 | 0.78 | 0.85 | 0.95 | 0.68 | 0.61 |
Continuous coordination
With GSD, all three models obtain higher diplomatic consistency than with SAD, while temporal-consistency changes differ by model. Claude gains most in temporal consistency (1.22 → 2.33), GPT gains less (1.22 → 1.44), and Qwen stays low (0.44 → 0.33). GPT attains the highest diplomatic consistency but also communicates substantially less.
- Runs graded 0 or 1 on temporal consistency
- 92 / 108
- GPT diplomatic consistency, SAD → GSD (7 applicable GSD runs)
- 2.44 → 3.71
- GPT diplomatic use, SAD → GSD
- 3.11 → 1.56
Nine runs per configuration. Each metric uses its own rubric; the three grades are not combined.
| TCTemporal consistency | DCDiplomatic consistency | DUDiplomatic use | |
|---|---|---|---|
| Qwen | |||
| SAD | |||
| LMAD | |||
| MAD | |||
| GSD | |||
| GPT | |||
| SAD | |||
| LMAD | |||
| MAD | |||
| GSD | |||
| Claude | |||
| SAD | |||
| LMAD | |||
| MAD | |||
| GSD | |||
* GPT–GSD diplomatic consistency averages its 7 applicable runs; two runs are N/A. Grades are LLM-assisted and non-blinded.
Grade distributions (Figure 8)Process traces
Linked message, task, and state records show where a revision is produced, passed downstream, or lost. Continuous traces separate finding a correction from preserving its accepted version.
A removed target returns
An earlier final command replaces an aerial target that lacks support in the visible entity list with ground screening, and that revision reaches task management and control. A later draft again excludes the old target, but finalization receives both task versions and restores it. Across 26 MAD archives, 32 of 1,002 changed task occurrences return to the supplied old wording; these counts describe textual returns whose correctness depends on context.
One decision, three local revisions
In the draft, the same approach of two enemy GunCars triggers both defensive fire and withdrawal. After review, their approach permits fire, while withdrawal requires escalation by the artillery group or additional ground forces; the final command, task management, and the task supplied to control all retain this distinction. The same decision revises a planned target, and its final message requests a verifiable source for Green’s claim of six bombers while keeping cooperation tentative. Execution of both conditional branches remains unverified.
Key Findings
Additional deliberation addresses different difficulties at different starting points, and its gains depend on preserving useful assignments through synthesis, command generation, and later rounds.
The highest-scoring workflow varies by model
Every multi-role workflow raises the snapshot total, but by different amounts: Qwen obtains its highest mean with MAD (8.89 → 11.12), GPT with GSD (11.07 → 13.17), and Claude with MAD (14.50 → 15.11). Close differences need not establish superiority.
Added roles address different errors
In the traced cases, advisers can supply threats and communication actions omitted by Qwen, and specialist analysis and review can resolve conflicting rendezvous conditions, cooperation targets, and task dependencies in Claude’s plans. GSD can improve GPT’s capability matching and parallel allocation, while Qwen’s GSD trace keeps a rule-invalid final cancellation rationale.
Content gains can be lost at finalization
Claude’s MAD gains 1.27 points across the five content dimensions but loses 0.66 in command compliance; its GSD gains 1.10 content points and loses 1.01, leaving a total gain of only 0.09.
Checking a claim is not offering a deal
LMAD improves claim checking for all three models, while cooperation moves in different directions (+0.32 for Qwen, −0.31 for GPT, −0.44 for Claude): a plan can verify a counterpart’s claim yet fail to state what the faction offers in return.
A correction can be found and still be lost
Task continuity remains a common difficulty: 92 of 108 runs receive temporal-consistency grades of 0 or 1. A correction can be found and passed downstream, yet a later finalization can restore the older task.
Consistency can come with silence
GSD raises diplomatic consistency for all three models, but GPT–GSD’s diplomatic use falls from 3.11 to 1.56: it maintains commitment conditions relatively consistently while using communication less often to obtain missing information.
Implications for workflow design
Assign advisers and reviewers to specific errors
Specify first what a role should check — whether rules constrain actions, whether tasks compete for resources, or whether cooperation terms constrain military assignments — then inspect whether its corrections survive synthesis.
Validate predictions and preserve accepted state
Check predicted interactions against entity rules and the resulting task set for unnecessary loss of ongoing work, and explicitly distinguish the current accepted task from historical context during finalization.
Treat communication as information gathering
Preserve uncertainty when evidence is missing, formulate a concrete request that could resolve it, and specify how the answer would change tasks.
The comparisons characterize complete model–workflow configurations, which jointly vary role allocation, prompt contents, computation, and response opportunities. Stage traces offer traceable explanations for score differences; they do not separately identify the causal contribution of individual roles or checks. The proposed checks follow from the observed processes; their effectiveness has yet to be measured as interventions.
Code & Data
The platform implementation, experiment configurations, and evaluation materials associated with this study.
Project Code
Platform implementation and experiment configurations.
Experimental Results
Archived evaluation records and analysis materials.