Skip to content

Add MiniMax-M3 NVFP4 GB300 vLLM Disaggregated AgentX with EAGLE3-GQA MTP / 新增 MiniMax-M3 NVFP4 GB300 vLLM 分离式 AgentX EAGLE3-GQA MTP 配置 - #2663

Merged
Ankur-singh merged 7 commits into
mainfrom
minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg
Aug 21, 2026
Merged

Add MiniMax-M3 NVFP4 GB300 vLLM Disaggregated AgentX with EAGLE3-GQA MTP / 新增 MiniMax-M3 NVFP4 GB300 vLLM 分离式 AgentX EAGLE3-GQA MTP 配置#2663
Ankur-singh merged 7 commits into
mainfrom
minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg

Conversation

@hshrivastava-droid

@hshrivastava-droid hshrivastava-droid commented Aug 18, 2026

Copy link
Copy Markdown
Collaborator

Description

Add minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg: MiniMax-M3 NVFP4 GB300 disaggregated AgentX config.

新增 minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg:MiniMax-M3 NVFP4 GB300 分离式 AgentX 配置。

  • Config: minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg
  • Image: vllm/vllm-openai:nightly-5e35a6f4f9bbc217c599692157ca985c894373f7

Related Issue

N/A

Type of Change

  • New feature
  • Configuration change

Checklist

  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /reuse-sweep-run on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

Note

Medium Risk
Mostly additive benchmark recipes, but launch_gb300-nv.sh now remaps CONFIG_FILE when EVAL_ONLY is set and copies MiniMax-M3 recipes into srt-slurm, which can change GB300 agentic job selection.

Overview
Adds minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg: disaggregated AgentX on GB300 with Dynamo-vLLM, NIXL P2P, and Mooncake DRAM KV offload.

Five Pareto points (1p1d / 1p3d / 2p5d, TP/EP mixes, conc 1–120) each have a throughput recipe with synthetic EAGLE3-GQA AL 2.78 and a real-verification eval twin. Thinking is pinned on (dyn-default-thinking-mode plus AIPERF_EXTRA_INPUTS=thinking:true).

launch_gb300-nv.sh maps MiniMax-M3 FP4 to /scratch/models/MiniMax-M3-NVFP4, copies the new srt-slurm recipes, and selects EVAL_CONFIG_FILE when EVAL_ONLY=true. Changelog entry appended.

Reviewed by Cursor Bugbot for commit 9727eb0. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

Comment on lines +147 to +157
command: "bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh"
env:
INFMAX_CONTAINER_WORKSPACE: "/infmax-workspace"
RESULT_DIR: "/logs/agentic"
PORT: "8000"
IS_MULTINODE: "true"
AIPERF_HTTP_X_DYNAMO_SESSION_ID_FROM_CORRELATION_ID: "true"
AIPERF_DATASET_MMAP_CACHE_DIR: "/aiperf_mmap_cache"
AIPERF_SERVER_METRICS_COLLECTION_INTERVAL: "1.0"
HF_HUB_CACHE: "/hf_hub_cache"
WEKA_LOADER_OVERRIDE: "semianalysis_cc_traces_weka_062126"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 All 10 new MiniMax-M3 GB300 agentic recipes under gb300-fp4/ omit AIPERF_EXTRA_INPUTS: "thinking:true" from benchmark.env, unlike every existing MiniMax-M3 agentic recipe (e.g. the GB200 sibling disagg-1p1d-dep8-dep4-agentic.yaml line 113, and all agg-*-agentic.yaml files). Since MiniMax-M3 is a reasoning model and this flag is what tells aiperf to exercise thinking mode, these runs will silently benchmark with thinking disabled or defaulted, producing output-length/acceptance-length/latency numbers that are not comparable to the model's real agentic-coding behavior or its own GB200 curve; add the flag to all 10 recipes' benchmark.env blocks before merging.

Extended reasoning...

The bug. Every one of the 10 new recipes under benchmarks/multi_node/srt-slurm-recipes/vllm/minimax-m3/agentic/gb300-fp4/ (5 throughput + 5 eval variants) defines a benchmark.env block that ends at WEKA_LOADER_OVERRIDE and never sets AIPERF_EXTRA_INPUTS. Every pre-existing MiniMax-M3 agentic recipe does set it — confirmed by grep, AIPERF_EXTRA_INPUTS: "thinking:true" appears in all 6 files under minimax-m3/gb200-fp4/agentic/ (the 4 agg-*-agentic.yaml files and disagg-1p1d-dep8-dep4-agentic.yaml line 113), and nowhere else in the repo except benchmarks/benchmark_lib.sh (the consumer).

How the flag is used. benchmark_lib.sh:2052-2053 shows the mechanism directly:

if [ -n "${AIPERF_EXTRA_INPUTS:-}" ]; then
    REPLAY_CMD+=" --extra-inputs $AIPERF_EXTRA_INPUTS"
fi

This is a plain conditional with no fallback/default — if the recipe's benchmark.env doesn't set AIPERF_EXTRA_INPUTS, aiperf never receives --extra-inputs thinking:true and there is no other code path that injects it. This is exactly why every existing MiniMax-M3 recipe sets it explicitly: thinking mode is not on by default for this client-side aiperf flag.

Why the shared dataset doesn't save it. One might assume the WEKA_LOADER_OVERRIDE: "semianalysis_cc_traces_weka_062126" dataset implicitly drives thinking behavior, but the direct GB200 sibling recipe (disagg-1p1d-dep8-dep4-agentic.yaml) uses the exact same dataset value AND still sets thinking:true explicitly — so the dataset alone doesn't enable it. reasoning-parser: "minimax_m3" (present in both GB200 and GB300 recipes) is a separate, server-side vLLM response-parsing setting; it does not make aiperf send requests requesting thinking mode.

Why this looks like an authoring accident rather than an intentional change. The new GB300 benchmark.env blocks are structurally different from the GB200 ones — they drop AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING, AIPERF_DYNAMO_SESSION_TIMEOUT_SECONDS, and AIPERF_REQUIRED_SERVER_METRIC_PREFIX, while adding AIPERF_SERVER_METRICS_COLLECTION_INTERVAL. This is consistent with the recipes having been templated from the DeepSeek-V4 agentic vLLM recipes (a non-reasoning model family that never sets this flag) rather than from the MiniMax-M3 GB200 agentic pattern. Nothing in the PR description documents disabling thinking mode as an intended change.

Impact. These recipes exist specifically to produce a GB300 curve comparable to the GB200 MiniMax-M3 agentic curve, and to power a "real-verification" eval twin plus a tuned synthetic_acceptance_length: 2.78 for the thinking workload. Benchmarking with thinking silently off means: (1) the throughput numbers (output length, acceptance length, latency) reflect a different workload than intended and aren't comparable to the GB200 sibling; (2) the eval twin verifies accuracy under the wrong mode; (3) the synthetic acceptance length tuning (2.78) was presumably chosen for the thinking workload and may not apply.

Proof walkthrough.

  1. Recipe gb300-fp4/1p1d-tep4-tp4-c1-agentic.yaml sets benchmark.env (lines 147-157) with WEKA_LOADER_OVERRIDE as the last key — no AIPERF_EXTRA_INPUTS.
  2. runners/launch_gb300-nv.sh exports every benchmark.env key as an environment variable for the agentic benchmark job (bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh).
  3. Inside that flow, benchmark_lib.sh:2052 checks [ -n "${AIPERF_EXTRA_INPUTS:-}" ] — since the var was never exported, this is false, and --extra-inputs thinking:true is never appended to REPLAY_CMD.
  4. aiperf therefore issues requests without the thinking-mode extra input, while the GB200 sibling recipe (same dataset, same model) does append it and does exercise thinking mode.
  5. Result: the GB300 run's output-length/acceptance-length/latency distribution reflects non-thinking (or default) behavior, not comparable to the GB200 curve it's meant to sit alongside.

Fix. Add AIPERF_EXTRA_INPUTS: "thinking:true" to the benchmark.env block of all 10 new gb300-fp4 recipe files, matching the GB200 pattern.

Comment on lines +76 to +84
served-model-name: "nvidia/MiniMax-M3-NVFP4"
kv-cache-dtype: "fp8"
block-size: 128
trust-remote-code: true
enable-prefix-caching: true
language-model-only: true
no-enable-flashinfer-autotune: true
reasoning-parser: "minimax_m3"
max-cudagraph-capture-size: 512

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 All 10 new gb300-fp4 MiniMax-M3 agentic/eval-agentic recipes set only reasoning-parser: minimax_m3 in prefill and decode, but omit dyn-tool-call-parser: minimax_m3 and dyn-reasoning-parser: minimax_m3 — args every existing MiniMax-M3 dynamo+vllm agentic recipe sets (e.g. gb200-fp4/agentic/disagg-1p1d-dep8-dep4-agentic.yaml lines 71-73, 97-99). Without these, the Dynamo frontend falls back to default tool-call/reasoning parsing for an agentic-coding workload that depends on correct tool-call extraction, which will misparse MiniMax-M3 output (especially in the real-verification -eval twins). Please add both dyn-tool-call-parser: "minimax_m3" and dyn-reasoning-parser: "minimax_m3" to the prefill and decode vllm_config blocks in all 10 new files.

Extended reasoning...

What the bug is: The 10 new benchmarks/multi_node/srt-slurm-recipes/vllm/minimax-m3/agentic/gb300-fp4/*.yaml recipes configure vLLM's engine-side reasoning-parser: "minimax_m3" in both the prefill and decode vllm_config blocks, but never set dyn-tool-call-parser or dyn-reasoning-parser. These two dyn-* keys are distinct, Dynamo-frontend-side directives — they tell the Dynamo frontend (not the vLLM engine) how to extract tool-call and reasoning content out of the raw model output stream before it's returned to the client/benchmark harness. reasoning-parser alone only configures vLLM's own internal parsing; it does not populate the Dynamo frontend's extraction path.\n\nWhere this diverges from precedent: Every existing MiniMax-M3 dynamo+vllm agentic recipe in this repo sets all three keys together, in both prefill and decode blocks. For example, gb200-fp4/agentic/disagg-1p1d-dep8-dep4-agentic.yaml sets reasoning-parser, dyn-tool-call-parser, and dyn-reasoning-parser, all to "minimax_m3", at lines 71-73 (prefill) and 97-99 (decode). Grepping all six gb200-fp4/agentic/*.yaml files (the agg-* and disagg-* siblings) confirms the same triad in every one. Grepping the 10 new gb300-fp4 files for dyn-tool-call-parser/dyn-reasoning-parser returns zero matches — only reasoning-parser (at line 83 prefill / line 103 decode in the representative 1p1d-tep4-tp4-c1-agentic.yaml) is present.\n\nWhy nothing else catches this: There's no schema validation on these recipe YAMLs enforcing that a given reasoning-parser value must be paired with matching dyn-* frontend args — the pairing is purely a convention followed by hand in every prior recipe. Nothing in launch_gb300-nv.sh or the srt-slurm framework injects a default derived from reasoning-parser; omitting the dyn-* keys just means the frontend uses whatever its built-in default parser is (likely none, or a generic one), not MiniMax-M3's actual tool-call format.\n\nImpact: These recipes exist specifically to benchmark an agentic-coding workload, where the benchmark harness needs to correctly extract tool calls from the model's output to measure/verify tool use. Without dyn-tool-call-parser: minimax_m3, the Dynamo frontend will fail to correctly delimit and extract MiniMax-M3's tool-call blocks, corrupting the agentic benchmark's traces. This is doubly important for the five *-eval-agentic.yaml twins, which run real MTP verification and rely on correctly-parsed tool calls to score correctness — a frontend parsing mismatch there would silently produce wrong eval results rather than an obvious crash.\n\nStep-by-step proof:\n1. Open gb200-fp4/agentic/disagg-1p1d-dep8-dep4-agentic.yaml (an existing, presumably-correct MiniMax-M3 dynamo+vllm agentic recipe) and look at the prefill vllm_config block: it sets reasoning-parser: "minimax_m3", dyn-tool-call-parser: "minimax_m3", and dyn-reasoning-parser: "minimax_m3" together (lines 71-73); the decode block repeats the same triad (lines 97-99).\n2. Now open the new 1p1d-tep4-tp4-c1-agentic.yaml added by this PR: the prefill block (around line 83) has only reasoning-parser: "minimax_m3"; the decode block (around line 103) has only the same single key. dyn-tool-call-parser and dyn-reasoning-parser do not appear anywhere in the file.\n3. Repeating this diff across the other 9 new gb300-fp4 files (both agentic and eval-agentic twins for each of the 5 topologies) shows the identical pattern: zero occurrences of either dyn-* key.\n4. Since reasoning-parser is vLLM-engine-scoped and the dyn-* keys are Dynamo-frontend-scoped (as evidenced by every other MiniMax-M3 recipe setting them independently and together), the frontend in these 10 new recipes has no MiniMax-M3-specific tool-call/reasoning extraction configured and falls back to its default behavior — which will misparse this model's actual output format.\n\nFix: Add dyn-tool-call-parser: "minimax_m3" and dyn-reasoning-parser: "minimax_m3" alongside reasoning-parser: "minimax_m3" in both the prefill and decode vllm_config blocks of all 10 new files, matching the established pattern in the gb200-fp4 agentic recipes.

Comment thread perf-changelog.yaml Outdated
@github-actions

Copy link
Copy Markdown
Contributor

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the previously flagged findings on this PR, this run also checked and ruled out three additional candidates: the new gb300-fp4 recipes omitting max-model-len (consistent with the existing GB200 MiniMax-M3 agentic recipes, which also omit it), the decode-side kv-transfer-config using kv_role: kv_both at the top level (the sub-connector that actually consumes, MooncakeStoreConnector, is explicitly set to kv_consumer; the top-level role only governs the MultiConnector/NixlConnector pairing), and the omission of a cudagraph_mode compilation-config key (not set in any other MiniMax-M3 agentic recipe either).

Extended reasoning...

Checked three candidate issues raised by finder agents against the existing MiniMax-M3 GB200 agentic recipes and the recipe'''s own kv-transfer-config structure: none diverge from established precedent in this repo, so none were escalated as bugs. This is separate from the two previously-posted findings on this PR (missing AIPERF_EXTRA_INPUTS thinking flag, missing dyn-tool-call-parser/dyn-reasoning-parser), which remain as inline comments and are unaffected by this note.

@github-actions

Copy link
Copy Markdown
Contributor

@Oseltamivir Oseltamivir left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

Oseltamivir and others added 3 commits August 21, 2026 10:17
合并 main,并解决 PR #2663 中的配置、性能变更日志和 GB300 启动脚本冲突。
…AL / 修复 MiniMax-M3 GB300 AgentX 思考模式以匹配黄金 AL

The 10 GB300 disagg recipes pin synthetic_acceptance_length 2.78, which is
golden_al_distribution/minimaxm3_eagle3_gqa.yaml minimax-m3.thinking_on[3],
but nothing enabled thinking, so the sweep served thinking-off traffic
(thinking_off[3] = 2.97). Eval logs from run 32194995479 confirm it: ~99.4%
of gsm8k responses opened with a bare </mm:think> empty reasoning block.

Adopt the GB200 MiniMax-M3 AgentX pattern, which is the established shape for
this dynamo-vllm agentic stack:
  - AIPERF_EXTRA_INPUTS: thinking:true in benchmark.env
  - dyn-tool-call-parser / dyn-reasoning-parser alongside reasoning-parser
    in both the prefill and decode vLLM configs
All 6 merged gb200-fp4/agentic recipes carry all three; these 10 carried none.

Requires a fresh full sweep: acceptance target and token profile both change.
…yn-default-thinking-mode 覆盖评测配置

AIPERF_EXTRA_INPUTS only reaches the aiperf replay. The eval twin runs
lm_eval, whose payload srt-slurm builds from a fixed env allow-list
(cli/do_sweep.py), so benchmark.env can never reach it. Left alone, the
accuracy gate would keep validating a different thinking mode than the
throughput arm it is meant to certify.

--dyn-default-thinking-mode is absorbed by dynamo.vllm's parse_known_args
before the strict AsyncEngineArgs parse, published to the runtime config
for non-prefill workers only (vllm/main.py:699-707), and applied by the
frontend to any chat request that carries no explicit thinking control.
Both arms hit the same frontend, so decode-side is enough to cover them.

Set on the decode block only; prefill workers never publish it. The flag
landed in dynamo 32c8091a8, before the 1.4.0.dev20260730 nightly cut this
config pins, and is absent from the 1.3.1 the GB200 siblings use, which is
why they had no server-side option available.
@github-actions

Copy link
Copy Markdown
Contributor

@Ankur-singh

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 32505173269

@Ankur-singh

Copy link
Copy Markdown
Collaborator

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this. — https://github.com/SemiAnalysisAI/InferenceX/actions/runs/32505173269
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this. — https://github.com/SemiAnalysisAI/InferenceX/actions/runs/32505173269
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in golden_al_distribution/ for that model, thinking mode, and draft length.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this.
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo.
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.).
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • N/A — this PR adds no single-node recipes. See the additional detail section.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped.
  • If this PR uses append-only: true — N/A, the new changelog entry does not use append-only.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.

Additional detail section:

Disclosure first. Two commits on this branch are mine — e4b0828d6 and d30a3697a, the thinking-mode fix described below. I am signing off on a PR I partly contributed to, so I have leaned on measured evidence from the run artifacts rather than on my own assertion that the change is correct. A second core-maintainer eye on those two commits is welcome.

Recipe link is N/A. This is a multi-node, disaggregated submission: all ten benchmark files are under benchmarks/multi_node/srt-slurm-recipes/vllm/minimax-m3/agentic/gb300-fp4/, and the master entry minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg sets multinode: true and disagg: true with framework: dynamo-vllm. The upstream recipe-link requirement covers single-node serve commands only.

Image. vllm/vllm-openai:nightly-5e35a6f4f9bbc217c599692157ca985c894373f7 — upstream vLLM Docker Hub org, running as shipped. No .patch, git apply, sed -i on engine sources, site-packages edit, monkey-patching, or forked engine wheel anywhere in the diff, so no waiver is required. This is also the first minimaxm3 entry on cluster:gb300-nv; dynamo-vllm runs the upstream vLLM engine itself with Dynamo as router/frontend, so the engine-first ordering guideline is satisfied rather than bypassed.

Golden AL, and the thinking-mode correction. The recipes pin synthetic_acceptance_length: 2.78 with num_speculative_tokens: 3 against the Inferact/MiniMax-M3-EAGLE3-GQA draft head, which is exactly golden_al_distribution/minimaxm3_eagle3_gqa.yamlminimax-m3.thinking_on[3] = 2.78. Note that is the GQA curve; the non-GQA minimaxm3_eagle3.yaml reads 2.83 at the same draft length and is correctly not the one used.

As originally submitted this config pinned that thinking_on acceptance target without ever enabling thinking. MiniMax-M3's chat template falls through to its adaptive branch when thinking_mode is undefined, so the model decided per-request and largely did not think. The evidence was in the first sweep (run 32194995479): 10484 of 10552 gsm8k responses — 99.36% — opened with a bare </mm:think>, an empty reasoning block. That is neither the thinking_on column nor the thinking_off one (2.97).

Commits e4b0828d6 and d30a3697a fix that by adopting the GB200 MiniMax-M3 AgentX shape plus a server-side default: AIPERF_EXTRA_INPUTS: "thinking:true" in benchmark.env, dyn-tool-call-parser and dyn-reasoning-parser on both prefill and decode, and dyn-default-thinking-mode: "enabled" on the decode worker only (prefill workers never publish it, per dynamo/vllm/main.py:699-707).

I verified the fix took effect from the run artifacts rather than assuming it. The obvious </mm:think> statistic goes 99.36% → 0.00%, but that measurement is confounded: dyn-reasoning-parser now routes reasoning into reasoning_content, which lm_eval discards, so it would read 0% whether thinking were on or off. The load-bearing evidence is the server-side generated-token count on the eval arm, integrated from the decode workers' vLLM logs over the eval window (1319 gsm8k requests each):

eval config old tokens/req new tokens/req ratio
1P (TP2) x 1D (TP4) c24 99.2 225.1 2.27x
2P (TP2) x 5D (TP2) c120 108.2 222.2 2.05x

Visible content is ~65-70 tokens in both runs, so the new run emits roughly 125-155 additional hidden tokens per request — a reasoning block. The eval arm carries no client-side thinking control at all (lm_eval's command line is byte-identical between the two runs), so that increase can only have come from the decode worker's dyn-default-thinking-mode. That also answers a question I could not settle from source: in this 1-prefill / N-decode disaggregated topology, the Dynamo frontend does resolve the decode worker's runtime config, so the server-side default is not a no-op on the eval arm. Supporting signals: 98.5% of generations differ from the old run, and prefill prompt tokens shift by -2.7/request, i.e. the rendered chat template genuinely changed.

What this fix does not do is move the published numbers, and I want that on the record. The AgentX replay forces output length from the trace — osl_mismatch_diff_pct == 0.0 for 100% of profiled requests in both runs, with identical sum(OSL) — and acceptance is synthetic at 2.78 regardless. Per-GPU total throughput moved 29,006 → 29,110 tps, which is noise. The fix is therefore about methodological consistency: the served workload now actually matches the thinking mode the pinned acceptance length was measured under, which is what the checklist requires. It is not a correction of inflated throughput.

Evals. gsm8k em_strict on run 32505173269, all four topologies: 0.96967 (1P x 1D c24), 0.97271 (1P x 3D c48), 0.97195 (2P x 5D c120), 0.97119 (1P TP1/EP4/DPA x 3D c24). utils/evals/thresholds.yaml has no minimaxm3 override, so the default: gsm8k: 0.90 bar applies and all four clear it comfortably. For comparison the pre-fix run scored 0.94845-0.96058; the range tightened from a 0.012 span to 0.003, consistent with thinking being on. All eval jobs ran the same image as the config, and the eval twin recipes deliberately drop the synthetic-acceptance knobs so evals keep real target verification.

Sweep. https://github.com/SemiAnalysisAI/InferenceX/actions/runs/32505173269 on head e78bc8bd6, conclusion success: 6 executed multi-node agentic / jobs and 4 executed multi-node agentic eval / jobs, all green. /reuse-sweep-run 32505173269 is pinned above. Note the c1 operating point has no eval job by design — utils/matrix_logic/generate_sweep_configs.py:25 sets MIN_EVAL_CONC = 16 for multi-node agentic evals, documented at docs/eval-agentx-procedures.md:25 and covered by tests.

Changelog. Appended at the physical end; no historical entry edited.

Outstanding, and not resolved by this sign-off: the PR is currently CONFLICTING on perf-changelog.yaml against main. Per .github/AGENT_OPERATIONS.md the resolution is to merge origin/main, restore the changelog byte-for-byte from origin/main, then append only this PR's entry at the tail — never a 3-way merge. That will move the head SHA and this sign-off will need re-posting against the new commit; the pinned run stays valid because e78bc8bd6 remains in the PR's commit list.

One residual caveat. AIPERF_EXTRA_INPUTS: "thinking:true" and dyn-default-thinking-mode were added together, so this run cannot separate their contributions on the replay arm. The prompt-token shift there (-1 to -6 per turn) is the same magnitude as the eval arm's (-2.7), which the server-side default alone plausibly explains, so the client-side field may be redundant. It is harmless — aiperf reported records_error_dropped: 0 and no error categories — but this run does not independently prove Dynamo honours a top-level thinking field. It matches the merged GB200 siblings, so I have kept it.

Signed: ankur-singh

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

✅✅✅ Verdict: PASS ✅✅✅

✅ Check 0 (CODEOWNER): PASS — Ankur-singh is a named owner of configs/nvidia-master.yaml (last-matching line); all other changed paths carry only the * catch-all, covered by a recognized CODEOWNER.
✅ Check 1 (sweep on in-PR commit): PASS — head e78bc8bd6 (unchanged since sign-off) carries 6 executed multi-node agentic / + 4 multi-node agentic eval / check-runs, all success, from run 32505173269.
✅ Check 2 (evals pass): PASS — agg_eval_all.json from that run: gsm8k em_strict 0.96967 / 0.97271 / 0.97195 / 0.97119 across all four topologies (n_eff=1319 each), all above the default: gsm8k: 0.90 bar (no minimaxm3 override), on the PR's image vllm/vllm-openai:nightly-5e35a6f4f9bbc....
➖ Check 3 (recipe link): N/A — disaggregated/multi-node submission (all recipes under benchmarks/multi_node/srt-slurm-recipes/**, master entry multinode: true, disagg: true); the recipe-link requirement applies to single-node recipes only.
✅ Check 4 (reuse command): PASS — /reuse-sweep-run 32505173269 posted by Ankur-singh (COLLABORATOR).
✅ Check 5 (latest checklist): PASS — every item in the current docs/PR_REVIEW_CHECKLIST.md template is present and checked; the single-node recipe sub-item is marked N/A with justification in the additional detail section.
✅ Check 6 (upstream image / engine-first): PASS — image is upstream vllm/vllm-openai:nightly-.... First minimaxm3 entry on cluster:gb300-nv, but the exception is documented in the sign-off: dynamo-vllm runs the upstream vLLM engine itself (Dynamo as router/frontend), matching the merged GB200 minimaxm3-fp4-gb200-dynamo-vllm-* precedent and the MODELS.md PoR (native/upstream vLLM engine).
✅ Check 7 (no deprecated models): PASS — MiniMax-M3 agentic-coding with EAGLE3 is the active PoR arm per MODELS.md as of 2026-08-21.
✅ Check 8 (no architecture hacks): PASS — no --hf-overrides/model-config edits anywhere in the diff; kv-cache fp8 on an NVFP4 model is precision, not FLOPs removal, and evals pass.
✅ Check 9 (spec-decode via chat template): PASS — agentic replay drives /v1/chat/completions with --endpoint-type chat (benchmarks/benchmark_lib.sh:2000).
✅ Check 10 (no engine patches): PASS — no .patch/git apply/sed -i/heredoc rewrites/site-packages edits/forked engine wheels in the diff; the declared dynamo framework install is the benchmarked frontend, not an engine modification.
✅ Check 11 (golden AL): PASS — all four throughput recipes pin rejection_sample_method: synthetic + synthetic_acceptance_length: 2.78 on prefill and decode, matching golden_al_distribution/minimaxm3_eagle3_gqa.yamlminimax-m3.thinking_on[3] = 2.78 for the Inferact/MiniMax-M3-EAGLE3-GQA draft with num_speculative_tokens: 3, and the config genuinely runs thinking-on (dyn-default-thinking-mode: enabled + AIPERF_EXTRA_INPUTS: thinking:true). Eval twins correctly drop the synthetic knobs for real verification.
➖ Check 12 (append-only): N/A — the new perf-changelog.yaml entry does not use append-only: true.

Ankur-singh
Ankur-singh approved these changes Aug 21, 2026
…ynamo-vllm-agentic-mtp-disagg

# Conflicts:
#	perf-changelog.yaml
@Klaud-Cold

Copy link
Copy Markdown
Collaborator

✅✅✅ Verdict: PASS ✅✅✅

✅ Check 0 (CODEOWNER): PASS — Ankur-singh is a named owner of configs/nvidia-master.yaml; all other changed paths fall to the catch-all, which a recognized CODEOWNER satisfies.
✅ Check 1 (Passing sweep on in-PR commit): PASS — run 32505173269 on in-PR commit e78bc8bd6: 6 executed multi-node agentic / and 4 executed multi-node agentic eval / check-runs, all success (c1/c20 points have no eval by design: MIN_EVAL_CONC = 16, and c20 shares the c24 recipe's eval twin).
✅ Check 2 (Evals pass): PASS — downloaded agg_eval_all.json from run 32505173269: gsm8k em_strict 0.96967–0.97271 across all four topologies (n_eff 1319 each), above the default: gsm8k: 0.90 bar (no minimaxm3 override), on the same vllm/vllm-openai:nightly-5e35a6f4… image as this PR's config.
➖ Check 3 (Recipe link): N/A — disaggregated/multi-node submission (all benchmark files under benchmarks/multi_node/srt-slurm-recipes/**, master entry sets multinode: true, disagg: true, framework: dynamo-vllm); the recipe-link requirement applies to single-node recipes only.
✅ Check 4 (Reuse command): PASS — /reuse-sweep-run 32505173269 posted by Ankur-singh (COLLABORATOR).
✅ Check 5 (Latest checklist): PASS — every item in the current docs/PR_REVIEW_CHECKLIST.md template is present and checked; the single-node-recipe sub-item is an explained N/A.
✅ Check 6 (Upstream image / engine-first): PASS — image is upstream vllm/vllm-openai:nightly-5e35a6f4…; dynamo-vllm serves via the upstream vLLM engine with Dynamo as router/frontend (documented in the sign-off), matching the merged GB200 minimaxm3 dynamo-vllm precedent; first minimaxm3 entry on cluster:gb300-nv.
✅ Check 7 (Deprecated models): PASS — MODELS.md keeps MiniMax-M3 active for agentic coding and deprecates the non-EAGLE3 arms in favor of EAGLE3; this PR is the EAGLE3 arm.
✅ Check 8 (No architecture hacks): PASS — no --hf-overrides/model-file edits; language-model-only skips no computed FLOPs on this text-only workload; kv-cache fp8 is precision-only with passing evals.
✅ Check 9 (Spec-decode chat templates): PASS — AgentX replay drives the chat-completions endpoint with the model's chat template (AIPERF_EXTRA_INPUTS: "thinking:true" is a chat-completions body field).
✅ Check 10 (No engine patches): PASS — no .patch/git apply/sed -i/site-packages edits/forked wheels in the diff; launcher changes are harness-side, and the Dynamo install is the declared framework's router/frontend as in existing dynamo-vllm entries.
✅ Check 11 (Golden AL for agentic spec-decode): PASS — all throughput recipes pin "rejection_sample_method":"synthetic","synthetic_acceptance_length":2.78 with num_speculative_tokens: 3 on the Inferact/MiniMax-M3-EAGLE3-GQA draft, matching golden_al_distribution/minimaxm3_eagle3_gqa.yaml thinking_on[3] = 2.78, and thinking is pinned on (dyn-default-thinking-mode: enabled); eval twins correctly drop the synthetic knobs to keep real verification, and no synthetic knobs appear on non-agentic configs.
➖ Check 12 (Append-only): N/A — the new perf-changelog.yaml entry does not use append-only: true.

Note: the sign-off was posted against e78bc8bd6; head 9727eb0 (the changelog conflict resolution it announced) only appends this PR's changelog entry at the tail, and the pinned sweep commit remains in the PR.

@Ankur-singh
Ankur-singh merged commit 49ac5c0 into main Aug 21, 2026
30 checks passed
@Ankur-singh
Ankur-singh deleted the minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg branch August 21, 2026 22:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

5 participants