-
Notifications
You must be signed in to change notification settings - Fork 280
Pull requests: SemiAnalysisAI/InferenceX
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[AMD] [AGENTX] Kimi Perf Tuning
agentx
AgentX benchmarks, recipes, and infrastructure
AMD
full-sweep-fail-fast
#2795
opened Sep 1, 2026 by
ajith-sirra-amd
Collaborator
Loading…
[AMD][MI35X] Serve Qwen3.5 MXFP4-AttnFP8-V2 on the MI355X SGLang
full-sweep-fail-fast
#2793
opened Sep 1, 2026 by
yichiche
Collaborator
Loading…
[Klaud Cold] [AMD] Bump DSV4 MI355X vLLM 8k/1k to the 2026-09-01 nightly / [Klaud Cold] [AMD] 将 DSV4 MI355X vLLM 8k/1k 更新至 2026-09-01 nightly
AMD
full-sweep-enabled
#2792
opened Sep 1, 2026 by
jiacao-amd
Collaborator
Loading…
config(dsv4): set H200 CUDA graph capture sizes / 配置(dsv4):设置 H200 CUDA 图捕获尺寸
full-sweep-enabled
#2791
opened Sep 1, 2026 by
RohitNagraj
Collaborator
Loading…
4 of 9 tasks
perf(qwen3.5-fp4-b200-sglang-mtp): enable FlashInfer GDN for TP2/EP2 / 为 TP2/EP2 启用 FlashInfer GDN
full-sweep-enabled
#2790
opened Aug 31, 2026 by
Ankur-singh
Collaborator
Loading…
feat(glm52-agentx): add compact B300 Dynamo+TRT-LLM recipes / 新增紧凑布局的 B300 Dynamo+TRT-LLM 配方
full-sweep-fail-fast
#2789
opened Aug 31, 2026 by
Ankur-singh
Collaborator
Loading…
5 of 9 tasks
CollectiveX: wire-basis bandwidth, nccl-ep LL hold, routed HT window, AsyncLL, DSv4-Pro workload
#2786
opened Aug 30, 2026 by
Oseltamivir
Collaborator
Loading…
Add H100 MiniMax-M3 NVMe and tiered AgentX sweep
full-sweep-enabled
#2775
opened Aug 28, 2026 by
cquil11
Collaborator
Loading…
Refresh DeepSeek V4 GB300 TRT-LLM AgentX metrics
full-sweep-enabled
#2774
opened Aug 28, 2026 by
cquil11
Collaborator
Loading…
Refresh GLM-5.2 GB300 TRT-LLM AgentX metrics
full-sweep-enabled
#2773
opened Aug 28, 2026 by
cquil11
Collaborator
Loading…
Refresh MiniMax M3 B300 TRT-LLM AgentX metrics
full-sweep-enabled
#2772
opened Aug 28, 2026 by
cquil11
Collaborator
Loading…
Refresh MiniMax M3 B200 TRT-LLM AgentX metrics
full-sweep-enabled
#2771
opened Aug 28, 2026 by
cquil11
Collaborator
Loading…
Refresh Qwen3.5 GB300 TRT-LLM AgentX metrics
full-sweep-enabled
#2770
opened Aug 28, 2026 by
cquil11
Collaborator
Loading…
[AMD][Power] fix: wait for AMD telemetry to cover the benchmark window end before stopping the monitor / 修复 AMD 遥测在基准窗口结束前被截断导致功耗校验失败的问题
#2767
opened Aug 28, 2026 by
edwingao28
Collaborator
Loading…
Align CI business priority across node counts and model families
#2765
opened Aug 27, 2026 by
cquil11
Collaborator
Loading…
Update DSV4 B300 SGLang AgentX image and HiCache concurrency grid / 更新 DSV4 B300 SGLang AgentX 镜像和 HiCache 并发网格
full-sweep-enabled
#2759
opened Aug 27, 2026 by
yhyang201
Collaborator
Loading…
4 of 9 tasks
[Klaud Cold] qwen3.8next-fp8-mi355x-sglang-agentic: Qwen3.8-Flash-Next FP8 SGLang AgentX on MI355X / MI355X 上 Qwen3.8-Flash-Next FP8 SGLang AgentX 配方
full-sweep-fail-fast
#2754
opened Aug 27, 2026 by
functionstackx
Collaborator
Loading…
[Klaud Cold] qwen3.8next-fp4-b200-sglang-agentic-mtp: day-zero Qwen3.8-Flash-Next NVFP4 SGLang AgentX on B200 / B200 上 Qwen3.8-Flash-Next NVFP4 SGLang AgentX 首发配方
full-sweep-fail-fast
#2751
opened Aug 27, 2026 by
functionstackx
Collaborator
Loading…
[CI] Import the DCGM exporter via enroot registry syntax / 用 enroot registry 语法导入 DCGM exporter
#2735
opened Aug 26, 2026 by
edwingao28
Collaborator
Loading…
[Klaud Cold] minimaxm3-fp4-mi355x-atom-agentic-mtp: day-zero MiniMax-M3 MXFP4 ATOM AgentX recipe on MI355X / MI355X 上 MiniMax-M3 MXFP4 ATOM AgentX 首发配方
full-sweep-fail-fast
#2733
opened Aug 26, 2026 by
functionstackx
Collaborator
Loading…
test: validate upstream-native vLLM Router topologies
#2731
opened Aug 25, 2026 by
cquil11
Collaborator
Loading…
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.