Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions .github/workflows/test-repository-roadmap-docs.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
name: Validate repository roadmap

on:
pull_request:
paths:
- 'docs/repository_roles.md'
- 'docs/repository_roles_zh.md'
- 'tests/test_repository_roadmap_docs.py'
- '.github/workflows/test-repository-roadmap-docs.yml'
push:
branches: [main]
paths:
- 'docs/repository_roles.md'
- 'docs/repository_roles_zh.md'
- 'tests/test_repository_roadmap_docs.py'
- '.github/workflows/test-repository-roadmap-docs.yml'

permissions:
contents: read

jobs:
roadmap-contract:
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@v4
with:
ref: ${{ github.event.pull_request.head.sha || github.sha }}
persist-credentials: false
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python tests/test_repository_roadmap_docs.py -v
11 changes: 6 additions & 5 deletions docs/repository_roles.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,7 @@ The roadmap is a queue of testable outcomes, not a list reserved for maintainers
| [Realtime preview efficiency on L20-class GPUs](https://github.com/modelscope/FunASR/issues/3528) | After matching the number of partial messages, which refresh interval and partial window provide the best latency/throughput trade-off without silently skipping previews? | Client JSONL and `--log-decode-profile` server logs from the exact commit, with SPK, ping, audio, concurrency, partial window, and partial-message count held constant | Reproduction on L20, L4, A10, or other non-H100 GPUs; analysis of queue, encoder, and engine time |
| [AMD Windows Vulkan stability](https://github.com/modelscope/FunASR/issues/3479) | Does the current runtime reach model initialization and transcription on the reporter's AMD GPU, and where is the last successful initialization boundary if it does not? | Exact archive name and SHA256, GPU/driver/Windows versions, full initialization log, and a reporter hardware retest | AMD Windows hardware owners and Vulkan/llama.cpp contributors |
| [Complete public checkpoint functionality](https://github.com/modelscope/FunASR/issues/3496) | How should the missing CTC tensors be published from an authorized model-owner account and validated after upload? | Immutable model revision, file hashes, public clean-cache download, and real timestamp/diarization inference | Model owners with Hugging Face write access and checkpoint validation experience |
| [Upstream model integrations](https://github.com/huggingface/transformers/pull/46180) | Can Fun-ASR-Nano remain compatible with upstream Transformers while preserving pinned model-card and regression-test boundaries? | Exact-head upstream CI, focused local tests, model-card review, and maintainer review | Transformers reviewers and users who can test downstream loading before merge |
| [Native Transformers compatibility](https://github.com/QwenAudio/Fun-ASR) | Can the merged Fun-ASR-Nano integration retain downstream loading and transcription compatibility across supported Transformers versions? | Exact model and library revisions, focused transcription tests, and model-card capability boundaries | Transformers users who can reproduce downstream compatibility failures |

### Before claiming an item

Expand All @@ -112,19 +112,20 @@ Contributions and issue evidence may be written in Chinese or English. A roadmap

### Delivered

- **Fun-ASR-Nano native Transformers integration** - [huggingface/transformers#46180](https://github.com/huggingface/transformers/pull/46180) merged on 2026-09-09. Use the [native `-hf` checkpoint](https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-hf) and the [model repository's Transformers entry point](https://github.com/QwenAudio/Fun-ASR). The native generation path does not include the toolkit's auxiliary CTC branch; this milestone does not establish native timestamps, diarization, or realtime sessions.
- **Qwen3-ASR offline vLLM workflow** - [#3592](https://github.com/modelscope/FunASR/pull/3592) delivered the native `Qwen3ASRModel.LLM` example. The reporter confirmed completion and closed [#3419](https://github.com/modelscope/FunASR/issues/3419) on 2026-09-09. Its accuracy comparison used that report's audio, reference, and custom scorer, not a general CER guarantee.
- **Bounded realtime long-session state** — fixes merged via [#3214](https://github.com/modelscope/FunASR/pull/3214) and [QwenAudio/Fun-ASR#135](https://github.com/QwenAudio/Fun-ASR/pull/135), diagnostics shipped, and reporter evidence allowed [#3101](https://github.com/modelscope/FunASR/issues/3101) to close.
- **Stable application-facing APIs** — the toolkit now ships an OpenAI-compatible transcription server, health checks, browser and command-line smoke tests, and documented Python / CLI / HTTP / WebSocket entry points in the [deployment matrix](./deployment_matrix.md).
- **Industrial and edge deployment paths** — vLLM serving and signed release workflows are documented; the verified ten-platform [`runtime-llamacpp-v0.2.6`](https://github.com/modelscope/FunASR/releases/tag/runtime-llamacpp-v0.2.6) archives cover Linux, macOS, and Windows CPU/GPU variants, including a dedicated Windows CUDA architecture 120 package for RTX 50 / Blackwell GPUs.
- **Joint transcription and diarization** — the third-party [MOSS-Transcribe-Diarize](./moss_transcribe_diarize.md) model is available through `AutoModel` with local Transformers and vLLM backends, or as an independent native SGLang Omni service. SGLang Omni is not an `AutoModel` backend. The model produces timestamps and speaker labels in one pass without a separate external VAD or speaker model; OpenMOSS remains the model owner.
- **Repository roles and issue routing** — [#3203](https://github.com/modelscope/FunASR/issues/3203) tracks this document and the remaining model-weight and vLLM entry-point questions. It stays open until those questions have evidence and the reporter has time to confirm.
- **Repository roles and issue routing** - [#3203](https://github.com/modelscope/FunASR/issues/3203) tracks this document and contribution-entry feedback. The reporter has confirmed the vLLM A/B entry-point explanation; the broader roadmap issue remains open, and documentation delivery does not close it automatically.

### In progress

- **Fun-ASR-Nano native Transformers integration** — [huggingface/transformers#46180](https://github.com/huggingface/transformers/pull/46180) is in review; use the PR's exact-head CI and review state as the source of truth.
- **Native Transformers downstream compatibility** - keep the [native model entry point](https://github.com/QwenAudio/Fun-ASR) and downstream transcription regressions aligned with exact model and library revisions. The initial integration is delivered; compatibility work is not a pending initial merge.
- **Restore complete public checkpoint functionality** — [#3496](https://github.com/modelscope/FunASR/issues/3496) tracks missing CTC tensors needed by timestamp and diarization paths in the Hugging Face checkpoint.
- **Realtime preview efficiency and L20 validation** — [#3528](https://github.com/modelscope/FunASR/issues/3528) established that v1.3.9 appeared faster by silently skipping most partial previews while its event loop was blocked. The issue remains open for equal-work L20 profiling and a deliberate refresh/window policy; it is not treated as a resolved throughput regression.
- **Qwen3-ASR offline vLLM workflow** — [#3592](https://github.com/modelscope/FunASR/pull/3592) adds a tested native `Qwen3ASRModel.LLM` example. [#3419](https://github.com/modelscope/FunASR/issues/3419) remains open until the reporter's 8–9% CER result can be reproduced with an exact model revision, service configuration, and scoring script.
- **AMD Windows Vulkan validation** — [#3479](https://github.com/modelscope/FunASR/issues/3479) remains open for reporter hardware retesting against `runtime-llamacpp-v0.2.6`; publication of the archive is not evidence that the hardware crash is fixed.
- **AMD Windows Vulkan validation** - [#3479](https://github.com/modelscope/FunASR/issues/3479) remains open for the original affected hardware paths. A reporter verified successful RX 9070 XT execution with runtime v0.2.5 and AMD 26.8.1 together; the driver also changed, so this does not isolate the runtime fix or establish the original 7600M XT / 780M outcome. Published archives alone are not hardware acceptance evidence.

### Next

Expand Down
11 changes: 6 additions & 5 deletions docs/repository_roles_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@
| [L20 等 GPU 上的实时预览效率](https://github.com/modelscope/FunASR/issues/3528) | 对齐 partial 消息数量后,怎样选择刷新间隔和 partial window,才能在不静默跳过预览的前提下取得合适的延迟/吞吐平衡? | 基于 exact commit 的客户端 JSONL 和 `--log-decode-profile` 服务端日志;固定 SPK、ping、音频、并发数、partial window 与 partial 消息数量 | 在 L20、L4、A10 或其他非 H100 GPU 上复现,并分析 queue、encoder 与 engine 时间 |
| [AMD Windows Vulkan 稳定性](https://github.com/modelscope/FunASR/issues/3479) | 当前 runtime 能否在报告者的 AMD GPU 上完成模型初始化和转写;若不能,最后成功的初始化边界在哪里? | 精确压缩包名称和 SHA256、GPU/驱动/Windows 版本、完整初始化日志,以及报告者硬件复测 | AMD Windows 硬件所有者和 Vulkan/llama.cpp 贡献者 |
| [恢复公开 checkpoint 的完整能力](https://github.com/modelscope/FunASR/issues/3496) | 如何从有权限的模型所有者账号发布缺失 CTC tensors,并在上传后完成验证? | 不可变模型 revision、文件哈希、公开 clean-cache 回下载和真实时间戳/说话人推理 | 有 Hugging Face 写权限的模型所有者和 checkpoint 验证贡献者 |
| [上游模型集成](https://github.com/huggingface/transformers/pull/46180) | 如何让 Fun-ASR-Nano 保持 Transformers 上游兼容,同时保留固定的 model card 和回归测试边界? | exact-head 上游 CI、聚焦本地测试、model card review 与维护者 review | Transformers reviewer,以及能在合并前验证下游加载的用户 |
| [原生 Transformers 兼容性](https://github.com/QwenAudio/Fun-ASR) | 已合并的 Fun-ASR-Nano 集成能否在支持的 Transformers 版本中保持下游加载和转写兼容? | 精确模型与库版本、聚焦转写测试,以及 model card 的能力边界 | 能复现下游兼容性问题的 Transformers 用户 |

### 认领前

Expand All @@ -110,19 +110,20 @@

### 已交付

- **Fun-ASR-Nano 的 Transformers 原生集成** - [huggingface/transformers#46180](https://github.com/huggingface/transformers/pull/46180) 已于 2026-09-09 合并。请使用[原生 `-hf` checkpoint](https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-hf)与[模型仓库的 Transformers 入口](https://github.com/QwenAudio/Fun-ASR)。原生生成路径不包含工具包的辅助 CTC 分支;这项交付不代表已有原生时间戳、说话人区分或实时会话支持。
- **Qwen3-ASR 离线 vLLM 工作流** - [#3592](https://github.com/modelscope/FunASR/pull/3592) 已交付原生 `Qwen3ASRModel.LLM` 示例。报告者确认完成,并于 2026-09-09 关闭 [#3419](https://github.com/modelscope/FunASR/issues/3419)。其中的准确率对照基于该报告的音频、参考文本与自定义评分脚本,不是通用 CER 保证。
- **实时服务长会话状态有界** —— [#3214](https://github.com/modelscope/FunASR/pull/3214) 与 [QwenAudio/Fun-ASR#135](https://github.com/QwenAudio/Fun-ASR/pull/135) 已合并,诊断能力已发布,报告者证据使 [#3101](https://github.com/modelscope/FunASR/issues/3101) 可以关闭。
- **稳定的应用接口** —— 工具包现已提供 OpenAI-compatible 转写服务、健康检查、浏览器与命令行 smoke test,并在[部署矩阵](./deployment_matrix_zh.md)中列出 Python / CLI / HTTP / WebSocket 入口。
- **工业与边缘部署路径** —— vLLM 服务和签名发布流程已有文档;经验证的十平台 [`runtime-llamacpp-v0.2.6`](https://github.com/modelscope/FunASR/releases/tag/runtime-llamacpp-v0.2.6) 压缩包覆盖 Linux、macOS 与 Windows 的 CPU/GPU 变体,并提供面向 RTX 50 / Blackwell 的 Windows CUDA architecture 120 专用包。
- **联合转写与说话人识别** —— 第三方 [MOSS-Transcribe-Diarize](./moss_transcribe_diarize_zh.md) 模型已通过 `AutoModel` 接入本地 Transformers 与 vLLM 后端,也可通过原生 SGLang Omni 独立服务。SGLang Omni 不是 `AutoModel` backend。它在一次推理中生成时间戳与说话人标签,不需要额外的外部 VAD 或说话人模型;模型所有者仍是 OpenMOSS。
- **仓库职责与 issue 路由** —— [#3203](https://github.com/modelscope/FunASR/issues/3203) 继续跟踪本文档以及尚未回答完的模型权重和 vLLM 入口问题。在这些问题有证据且报告者有合理确认时间之前,issue 保持开放。
- **仓库职责与 issue 路由** - [#3203](https://github.com/modelscope/FunASR/issues/3203) 跟踪本文档与贡献入口反馈。报告者已确认 vLLM A/B 入口说明清楚;更广泛的路线图 issue 仍保持开放,不因文档交付而自动关闭。

### 进行中

- **Fun-ASR-Nano 的 Transformers 原生集成** —— [huggingface/transformers#46180](https://github.com/huggingface/transformers/pull/46180) 正在审查;以该 PR 的 exact-head CI 与 review 状态为准。
- **原生 Transformers 下游兼容性** - 按精确模型与库版本持续核对[原生模型入口](https://github.com/QwenAudio/Fun-ASR)与下游转写回归。初始集成已经交付;兼容性维护不代表首次合并仍在等待。
- **恢复公开 checkpoint 的完整能力** —— [#3496](https://github.com/modelscope/FunASR/issues/3496) 跟踪 Hugging Face checkpoint 缺少时间戳与说话人路径所需 CTC tensors 的问题。
- **实时预览效率与 L20 验证** —— [#3528](https://github.com/modelscope/FunASR/issues/3528) 已确认 v1.3.9 看似更快,是因为事件循环阻塞时静默跳过了大部分 partial 预览。该 issue 继续开放,用于等工作量的 L20 profiling 和明确的刷新/window 策略;不能把它当成已经解决的吞吐回退。
- **Qwen3-ASR 离线 vLLM 工作流** —— [#3592](https://github.com/modelscope/FunASR/pull/3592) 增加经过验证的原生 `Qwen3ASRModel.LLM` 示例。[#3419](https://github.com/modelscope/FunASR/issues/3419) 继续开放,直到能用精确模型 revision、服务配置和评分脚本复现报告者的 8–9% CER。
- **AMD Windows Vulkan 验证** —— [#3479](https://github.com/modelscope/FunASR/issues/3479) 保持开放,等待报告者在 `runtime-llamacpp-v0.2.6` 上进行硬件复测;发布压缩包不等于硬件崩溃已经修复。
- **AMD Windows Vulkan 验证** - [#3479](https://github.com/modelscope/FunASR/issues/3479) 保持开放,等待最初受影响的硬件路径确认。已有报告者在 runtime v0.2.5 与 AMD 26.8.1 的组合上验证 RX 9070 XT 成功运行;驱动也同时改变,因此不能单独归因于 runtime 修复,更不能据此确认原始 7600M XT / 780M 的结果。发布压缩包本身不等于硬件验收。

### 下一步

Expand Down
46 changes: 46 additions & 0 deletions tests/test_repository_roadmap_docs.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
"""Keep completed roadmap milestones separate from unresolved validation."""

from pathlib import Path
import unittest


ROOT = Path(__file__).resolve().parents[1]
GUIDES = (
("repository_roles.md", "### Delivered", "### In progress", "### Next"),
("repository_roles_zh.md", "### \u5df2\u4ea4\u4ed8", "### \u8fdb\u884c\u4e2d", "### \u4e0b\u4e00\u6b65"),
)


class RepositoryRoadmapContract(unittest.TestCase):
def sections(self, guide):
name, delivered, active, following = guide
text = (ROOT / "docs" / name).read_text(encoding="utf-8")
return (text.split(delivered, 1)[1].split(active, 1)[0],
text.split(active, 1)[1].split(following, 1)[0])

def test_native_transformers_merge_is_delivered(self):
for guide in GUIDES:
with self.subTest(guide=guide[0]):
delivered, active = self.sections(guide)
self.assertIn("huggingface/transformers/pull/46180", delivered)
self.assertIn("FunAudioLLM/Fun-ASR-Nano-2512-hf", delivered)
self.assertNotIn("huggingface/transformers/pull/46180", active)

def test_reporter_closed_qwen_workflow_is_delivered(self):
for guide in GUIDES:
with self.subTest(guide=guide[0]):
delivered, active = self.sections(guide)
self.assertIn("FunASR/pull/3592", delivered)
self.assertIn("FunASR/issues/3419", delivered)
self.assertNotIn("FunASR/issues/3419", active)

def test_unresolved_hardware_and_checkpoint_work_stays_open(self):
for guide in GUIDES:
with self.subTest(guide=guide[0]):
_, active = self.sections(guide)
for issue in (3496, 3528, 3479):
self.assertIn(f"FunASR/issues/{issue}", active)


if __name__ == "__main__":
unittest.main()
Loading