NVIDIAのNemotron 3 DiarizationをNeMo Speechで使ってみる

NVIDIAのNemotron 3 DiarizationをNeMo Speechで使ってみる

NVIDIAが最近リリースした新型話者識別モデル「Nemotron 3 Diarization」を実際に試してみました。モデルカードの手順では動作しなかったため、その解決方法も含めて紹介します。
2026.09.29

こんばんは、情報システム室の夏目です。

先日NVIDIAは話者識別用のDiarization ModelであるNemotoron 3 Diarizationをリリースしました。

https://huggingface.co/blog/nvidia/nemotron-diarization

VoiceArenaのDiarization-Benchリーダーボード で1位の成績で、最大8人の話者識別が可能です。

Hugging Faceのモデルカード で記載されていたNeMo Speechで試してみました。
しかし、記載されていた方法では使用できなかったのでどう使ったのかも記載します。

1. NeMo Speech

NeMo Speechは二つのものがあります。
NVIDIA NeMo Speech (Python) と NeMo-Speech.cpp (C++) です。

https://github.com/NVIDIA-NeMo/Speech

https://github.com/NVIDIA/NeMo-Speech.cpp

前者はNVIDIAの音声AI向けフレームワークで、学習やファインチューニングでのモデルの作成や推論で実際に動かすことができます。
nemo-toolkit としてPyPIで公開されています。

後者は、前者で作られたモデルをローカルで動かすためのCLIツールです。
ASR(音声書き起こし)やDiarization(話者分離)、TTS(text to speech)などを CPU/CUDA/Metal/Vulkan で動かすことができます。

2. 検証環境

今回試すにあたって以下の環境で試しました。

  • macOS Tahoe 26.6.2
  • uv (Python 3.14)

3. NVIDIA NeMo Speech (Python)で動かす

3-1. モデルカードを参考に動かしてみる

まずuvの環境を作成し、nemo-toolkitをインストールします。

$ uv init --python 3.14 --bare
Initialized project `trial-python`

$ uv add nemo-toolkit[asr]
Using CPython 3.14.7
Creating virtual environment at: .venv
Resolved 161 packages in 4.54s
Installed 140 packages in 787ms
 + absl-py==2.5.0
 + aiohappyeyeballs==2.7.1
 + aiohttp==3.14.3
 + aiosignal==1.4.0
 + aistore==1.26.0
 + annotated-doc==0.0.5
 + annotated-types==0.8.0
 + antlr4-python3-runtime==4.9.3
 + anyio==4.15.1
 + attrs==26.1.0
 + audioop-lts==0.2.2
 + audioread==3.1.0
 + braceexpand==0.1.7
 + certifi==2026.7.22
 + cffi==2.1.1
 + charset-normalizer==3.5.1
 + click==8.5.0
 + cloudpickle==3.1.2
 + colorama==0.4.6
 + cytoolz==1.1.0
 + datasets==5.0.1
 + decorator==5.3.1
 + dill==0.4.1
 + einops==0.8.2
 + filelock==4.0.3
 + frozenlist==1.8.0
 + fsspec==2025.12.0
 + googleapis-common-protos==1.75.4
 + grpcio==1.84.0
 + h11==0.16.0
 + hf-xet==1.6.0
 + httpcore==1.0.9
 + httpx==0.28.1
 + huggingface-hub==1.33.0
 + humanize==4.16.0
 + hydra-core==1.3.2
 + idna==3.20
 + indic-numtowords==1.1.0
 + intervaltree==3.2.1
 + jinja2==3.1.6
 + joblib==1.6.0
 + kaldialign==0.12.0
 + lazy-loader==0.6
 + lhotse==1.33.0
 + librosa==1.0.0
 + lightning==2.4.0
 + lightning-utilities==0.15.3
 + llvmlite==0.49.0
 + lxml==6.1.3
 + markdown==3.11
 + markdown-it-py==4.2.0
 + markupsafe==3.0.3
 + mdurl==0.1.2
 + ml-dtypes==0.6.0
 + more-itertools==11.1.0
 + mpmath==1.3.0
 + msgpack==1.2.2
 + msgspec==0.21.1
 + multidict==6.9.1
 + multiprocess==0.70.19
 + narwhals==2.26.0
 + nemo-toolkit==3.0.0
 + networkx==3.7
 + numba==0.67.0
 + numpy==2.5.3
 + nv-one-logger-core==2.3.1
 + nv-one-logger-pytorch-lightning-integration==2.3.1
 + nv-one-logger-training-telemetry==2.3.1
 + omegaconf==2.3.0
 + onnx==1.23.0
 + opentelemetry-api==1.45.0
 + opentelemetry-exporter-http-transport==0.66b0
 + opentelemetry-exporter-otlp-common==0.66b0
 + opentelemetry-exporter-otlp-proto-common==1.45.0
 + opentelemetry-exporter-otlp-proto-http==1.45.0
 + opentelemetry-proto==1.45.0
 + opentelemetry-sdk==1.45.0
 + opentelemetry-semantic-conventions==0.66b0
 + overrides==7.7.0
 + packaging==24.2
 + pandas==3.0.6
 + pillow==12.3.0
 + platformdirs==4.11.15
 + pooch==1.9.0
 + portalocker==4.4.0
 + propcache==0.5.4
 + protobuf==7.36.2
 + pyarrow==25.0.1
 + pycparser==3.0
 + pydantic==2.13.5
 + pydantic-core==2.46.5
 + pygments==2.21.0
 + python-dateutil==2.9.0.post0
 + pytorch-lightning==2.6.6
 + pyyaml==6.0.3
 + regex==2026.9.10
 + requests==2.34.2
 + rich==15.0.0
 + sacrebleu==2.6.0
 + safetensors==0.8.0
 + scikit-learn==1.9.1
 + scipy==1.18.1
 + sentencepiece==0.2.2
 + setuptools==84.0.0
 + shellingham==1.5.4
 + six==1.17.0
 + smart-open==8.0.1
 + sortedcontainers==2.4.0
 + soundfile==0.14.0
 + soxr==1.1.0
 + standard-aifc==3.13.0
 + standard-chunk==3.13.0
 + standard-sunau==3.13.0
 + strenum==0.4.15
 + sympy==1.14.0
 + tabulate==0.10.0
 + tenacity==9.1.4
 + tensorboard==2.21.0
 + tensorboard-data-server==0.7.2
 + text-unidecode==1.3
 + text2num==3.1.0
 + threadpoolctl==3.7.0
 + tokenizers==0.23.2
 + toml==0.10.2
 + toolz==1.1.0
 + torch==2.14.0
 + torchmetrics==1.9.0
 + tqdm==4.70.1
 + transformers==5.17.0
 + typer==0.27.2
 + typing-extensions==4.16.0
 + typing-inspection==0.4.4
 + urllib3==2.8.0
 + wandb==0.30.0
 + webdataset==1.0.2
 + werkzeug==3.1.8
 + whisper-normalizer==0.1.15
 + wrapt==2.4.1
 + xxhash==3.5.0
 + yarl==1.25.1

次にモデルカードに記載されていたコードを使ってみます。

main.py
from nemo.collections.asr.models import SortformerEncLabelModel
diar_model = SortformerEncLabelModel.from_pretrained("nvidia/Nemotron-3-Diarization")
diar_model.eval()

diar_model.sortformer_modules.chunk_len = 340
diar_model.sortformer_modules.chunk_right_context = 40
diar_model.sortformer_modules.fifo_len = 40
diar_model.sortformer_modules.spkcache_update_period = 300
diar_model._check_streaming_parameters()

predicted_segments = diar_model.diarize(audio=["audio.wav"], batch_size=1)

for segment in predicted_segments[0]:
    print(segment)
$ uv run main.py
    subsegment_nspk_bias: 1.5
    subsegment_start_guard_sec: 0.25
    subsegment_min_first_spk_sec: 0.5
    subsegment_splice_silence_sec: 0.1

[NeMo W 2026-09-29 15:45:15 modelPT:182] If you intend to do validation, please call the ModelP
    Validation config :
    manifest_filepath: null
    sample_rate: 16000
    num_spks: 8
    session_len_sec: -1
    soft_targets: false
    batch_size: 32
    shuffle: false
    num_workers: 18
    drop_last: false
    pin_memory: true
    use_lhotse: false

[NeMo W 2026-09-29 15:45:15 modelPT:189] Please call the ModelPT.setup_test_data() or ModelPT.s
    Test config :
    manifest_filepath: null
    sample_rate: 16000
    num_spks: 8
    session_len_sec: -1
    soft_targets: false
    batch_size: 32
    shuffle: false
    num_workers: 18
    drop_last: false
    pin_memory: true
    use_lhotse: false

[NeMo W 2026-09-29 15:45:15 experimental:26] `<class 'nemo.collections.asr.modules.transformer_
[NeMo E 2026-09-29 15:45:15 common:827] Model instantiation failed!
    Target class:       nemo.collections.asr.models.sortformer_diar_models.SortformerEncLabelMo
    Error(s):   Error in call to target 'nemo.collections.asr.modules.transformer_encoder.Trans
    ValueError("self_attention_model='rope' is not supported. Currently only 'abs_pos', 'rel_po
    Traceback (most recent call last):
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        return _target_(*args, **kwargs)
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        return wrapped(*args, **kwargs)
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        raise ValueError(
        ...<2 lines>...
        )
    ValueError: self_attention_model='rope' is not supported. Currently only 'abs_pos', 'rel_po

    The above exception was the direct cause of the following exception:

    Traceback (most recent call last):
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        instance = imported_cls(cfg=config, trainer=trainer)
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        self.encoder = SortformerEncLabelModel.from_config_dict(self._cfg.encoder).to(self.devi
                       ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        instance = safe_instantiate(config=config)
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        return hydra.utils.instantiate(config, *args, **kwargs)
               ~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        return instantiate_node(
            config, *args, recursive=_recursive_, convert=_convert_, partial=_partial_
        )
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        return _call_target(_target_, partial, args, kwargs, full_key)
      File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/
        raise InstantiationException(msg) from e
    hydra.errors.InstantiationException: Error in call to target 'nemo.collections.asr.modules.
    ValueError("self_attention_model='rope' is not supported. Currently only 'abs_pos', 'rel_po

Traceback (most recent call last):
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/
    return _target_(*args, **kwargs)
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/
    return wrapped(*args, **kwargs)
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/
    raise ValueError(
    ...<2 lines>...
    )
ValueError: self_attention_model='rope' is not supported. Currently only 'abs_pos', 'rel_pos',

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/main.py",
    diar_model = SortformerEncLabelModel.from_pretrained("nvidia/Nemotron-3-Diarization")
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/
    instance = class_.restore_from(
        restore_path=nemo_model_file_in_cache,
    ...<5 lines>...
        save_restore_connector=save_restore_connector,
    )
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/
    instance = cls._save_restore_connector.restore_from(
        cls,
    ...<6 lines>...
        validate_access_integrity,
    )
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/
    loaded_params = self.load_config_and_state_dict(
        calling_cls,
    ...<6 lines>...
        validate_access_integrity,
    )
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/nemo/core/connectors/save_restore_connector.py", line 193, in load_config_and_state_dict
    instance = calling_cls.from_config_dict(config=conf, trainer=trainer)
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/nemo/core/classes/common.py", line 828, in from_config_dict
    raise e
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/nemo/core/classes/common.py", line 820, in from_config_dict
    instance = cls(cfg=config, trainer=trainer)
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/nemo/collections/asr/models/sortformer_diar_models.py", line 123, in __init__
    self.encoder = SortformerEncLabelModel.from_config_dict(self._cfg.encoder).to(self.device)
                   ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/nemo/core/classes/common.py", line 782, in from_config_dict
    instance = safe_instantiate(config=config)
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/nemo/core/classes/common.py", line 348, in safe_instantiate
    return hydra.utils.instantiate(config, *args, **kwargs)
           ~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/hydra/_internal/instantiate/_instantiate2.py", line 226, in instantiate
    return instantiate_node(
        config, *args, recursive=_recursive_, convert=_convert_, partial=_partial_
    )
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/hydra/_internal/instantiate/_instantiate2.py", line 347, in instantiate_node
    return _call_target(_target_, partial, args, kwargs, full_key)
  File "/Users/natsume.yuta/spaces/work/blog/022_nemotron_3_diarization/trial_python/.venv/lib/python3.14/site-packages/hydra/_internal/instantiate/_instantiate2.py", line 97, in _call_target
    raise InstantiationException(msg) from e
hydra.errors.InstantiationException: Error in call to target 'nemo.collections.asr.modules.transformer_encoder.TransformerEncoder':
ValueError("self_attention_model='rope' is not supported. Currently only 'abs_pos', 'rel_pos', and 'no_pos' (or None) are available.")

ValueError("self_attention_model='rope' is not supported. Currently only 'abs_pos', 'rel_pos', and 'no_pos' (or None) are available.") というエラーが出て失敗しました。

3-2. なぜ失敗したのか

これはNemotoron 3 Diarizationで Rotary Positional Embeddings (RoPE) が採用されており、 nemo-toolkit[asr]の最新バージョン3.0.0 (8月上旬にリリース) では ropeに対応していないことが原因です。

現在のmainブランチではRoPE対応が入っているので、対応バージョンがリリースされるまでは、そちらを使えば動かすことができます。

3-3. nemo-toolkitをmainブランチのものに変えて使ってみる

pyproject.tomlを変更します。

pyproject.toml
[project]
name = "trial-python"
version = "0.1.0"
requires-python = ">=3.14"
dependencies = [
-   "nemo-toolkit[asr]>=3.0.0",
+   "nemo-toolkit[asr] @ git+https://github.com/NVIDIA/NeMo.git"
]

nemo-toolkitを再インストールします。

$ uv lock
    Updated https://github.com/NVIDIA/NeMo.git (cf724ac337d1ebc7d0dda1e23fb80916f52927a5)      Resolved 163 packages in 8.85s
Updated lhotse v1.33.0 -> v2.0.0a6
Updated nemo-toolkit v3.0.0 -> v3.1.0+cf724ac33 (cf724ac3)
Updated torch v2.14.0 -> v2.14.0, v2.14.0+cpu

$ uv sync
Resolved 163 packages in 16ms
Uninstalled 2 packages in 145ms
Installed 2 packages in 58ms
 - lhotse==1.33.0
 + lhotse==2.0.0a6
 - nemo-toolkit==3.0.0
 + nemo-toolkit==3.1.0+cf724ac33 (from git+https://github.com/NVIDIA/NeMo.git@cf724ac337d1ebc7d0dda1e23fb80916f52927a5)

再度実行してみます。

$ uv run main.py
Streaming Steps: 100%|████████████████████████████████████████| 47/47 [00:16<00:00,  2.92it/s]
Post-processing: 100%|██████████████████████████████████████████| 1/1 [00:00<00:00, 35.91it/s]
Diarizing: 1it [00:16, 16.68s/it]                                       | 0/1 [00:00<?, ?it/s]
5.000 7.020 speaker_0
7.340 10.120 speaker_0
10.330 14.580 speaker_0
14.810 17.440 speaker_0
17.890 19.130 speaker_0
32.150 32.960 speaker_0

... (中略) ...

1248.320 1251.610 speaker_0
1252.600 1255.430 speaker_0
1255.900 1259.210 speaker_0
1259.780 1260.570 speaker_0
1260.830 1263.970 speaker_0
19.480 23.480 speaker_1
23.760 27.390 speaker_1
27.760 31.920 speaker_1
63.150 63.410 speaker_1
63.670 69.710 speaker_1

... (中略) ...

1220.880 1224.490 speaker_1
1224.820 1228.830 speaker_1
1229.100 1231.280 speaker_1
1231.670 1234.700 speaker_1

無事に成功しました。

speaker_0が連続した後にspeaker_1が出ているので識別がうまくいっていないように見えます。
しかしよく秒数をよく見ると、話者ごとにまとめられているだけできちんと識別できてそうでした。

整形後(最初の15行)
5.000 7.020 speaker_0
7.340 10.120 speaker_0
10.330 14.580 speaker_0
14.810 17.440 speaker_0
17.890 19.130 speaker_0
19.480 23.480 speaker_1
23.760 27.390 speaker_1
27.760 31.920 speaker_1
32.150 32.960 speaker_0
33.240 36.550 speaker_0
36.820 44.020 speaker_0
44.470 46.640 speaker_0
46.910 54.770 speaker_0
55.040 62.800 speaker_0
63.150 63.410 speaker_1

4. NeMo-Speech.cpp (C++)を使ってみる

4-1. インストールする

GithubリポジトリのREADMEに記載されたインストールコマンドを実行します。

$ curl -fsSL https://github.com/NVIDIA/NeMo-Speech.cpp/raw/main/scripts/install.sh | sh
NeMo-Speech.cpp 0.1.0 (macos/aarch64, metal)
Artifact: https://github.com/NVIDIA/NeMo-Speech.cpp/releases/download/v0.1.0/nemo-speech-0.1.0-macos-aarch64-metal.tar.gz
Source:   https://github.com/NVIDIA/NeMo-Speech.cpp.git#v0.1.0 (metal-server fallback)
Prefix:   /Users/natsume.yuta/Library/Application Support/NeMoSpeech
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0
100 3383k  100 3383k    0     0  3152k      0  0:00:01  0:00:01 --:--:-- 3152k
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0
100   111  100   111    0     0    172      0 --:--:-- --:--:-- --:--:--   172
nemo-speech 0.1.0
Next: download a model, then run 'nemo-speech transcribe' or 'nemo-speech serve' (see README.md).

$ nemo-speech --help
NeMo-Speech.cpp 0.1.0

Usage: nemo-speech <command> [options]

Commands:
  transcribe   Transcribe an audio file, directory, or microphone
  diarize      Identify speaker segments in an audio file
  translate    Translate text
  synthesize   Synthesize speech to a WAV file
  bench        Benchmark an end-to-end ASR workload
  pull         Download a pinned model from Hugging Face
  model        List, pull, or inspect models
  doctor       Inspect runtime and device availability
  health       Check a running local HTTP server
  serve        Start the local API and playground
  help         Show help for a command

Global options:
  -h, --help       Show help
  --version        Show version
  --json           Emit machine-readable results and errors
  --quiet          Suppress non-result progress messages
  --verbose        Emit additional diagnostics on stderr

4-2. nemo-speech model list をしてみる

ヘルプを見るとmodelコマンドで使用できるモデルが見れそうなので少しみてみます

$ nemo-speech model --help
Usage: nemo-speech model <action> [arguments]

Actions:
  list             List indexed repositories and command defaults
  pull REPO        Download and verify an indexed Hugging Face repository
  info FILE        Inspect a local GGUF file

Models are cached under the platform user cache directory. Override it
with NEMO_SPEECH_MODEL_DIR. Downloads require curl on PATH.

$ nemo-speech model list
Available models (* = command default)

ASR — transcribe, bench, serve
  * nemotron-3.5
      repo: nvidia/nemotron-3.5-asr-streaming-0.6b
      also: nemotron-asr
    nemotron-en
      repo: nvidia/nemotron-speech-streaming-en-0.6b
    parakeet-ctc
      repo: nvidia/parakeet-ctc-1.1b
    parakeet-tdt
      repo: nvidia/parakeet-tdt-0.6b-v3

Diarization — diarize, transcribe --diarize, serve
  * sortformer
      repo: nvidia/diar_streaming_sortformer_4spk-v2
      also: sortformer-diar

TTS — synthesize, serve
  * magpie (speech model + tokenizer)
      repo: nvidia/magpie_tts_multilingual_357m
      also: magpie-tts
      pulls: nano-codec
  * nano-codec (codec)
      repo: nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
      also: nanocodec

Use a short name or full repository ID wherever MODEL is accepted.
Download ahead of time with: nemo-speech pull <name>

話者識別に使用するデフォルトのモデルが nvidia/diar_streaming_sortformer_4spk-v2 になっています。

モデルカードには nemo-speech diarize meeting.wav で使用できるとありましたが、デフォルトでは使えなさそうです。

4-3. Nemotron-3-Diarization.q8_0.gguf をダウンロードして使ってみる

モデルカードのページからファイル一覧に移動し、 /Users/natsume.yuta/Downloads/Nemotron-3-Diarization.q8_0.gguf をダウンロードして使ってみます。

https://huggingface.co/nvidia/Nemotron-3-Diarization

61350381-64e5-4282-a5f5-97fc361e4a57

9d1ed1af-a3e8-4f21-bab0-a2b5950fbaff

$ nemo-speech diarize -m Nemotron-3-Diarization.q8_0.gguf audio.wav
[nemo-speech] diarize session started
nemo-speech diarize: sortformer: pre_ln transformer variant is not supported
[nemo-speech] diarize session failed (exit code 1)

こちらも使えませんでした。

4-4. なぜダメだったのか

NeMo-Speech.cpp (C++)の最新版はv0.1.0で、Nemotron 3 Diarizationよりも前にリリースされたものです。
そのため、Nemotoron 3 Diarizationで採用されているアーキテクチャに対応していないようです。

こちらもmainブランチでは対応しているようです。

fa5df1ba-e82c-4053-b758-6e0b61e97b15

4-5. mainブランチからビルドして使ってみる

まずHomebrewでビルドに必要なものをインストールします。

$ xcode-select --install  # skip if the Command Line Tools are already installed
$ brew install cmake ninja sentencepiece abseil

次にgitからクローンし、そのディレクトリに入ります。

$ git clone git@github.com:NVIDIA/NeMo-Speech.cpp.git
$ cd NeMo-Speech.cpp

次に必要なサブモジュールをインストールします。

$ git submodule update --init ggml
$ git submodule update --init --recursive llama.cpp

次にビルドします。

$ scripts/configure.sh metal-speech
$ cmake --build --preset metal-speech

mac向けへのビルドなのでMetal向けのビルドとし、今回はASRやDiarizationなど向けのビルドも行うようにしています。

ビルドが成功すると build/metal-speech/bin/nemo-speech に実行ファイルが生成されます。

$ build/metal-speech/bin/nemo-speech --help
NeMo-Speech.cpp 0.1.0

Usage: build/metal-speech/bin/nemo-speech <command> [options]

Commands:
  transcribe   Transcribe an audio file, directory, or microphone
  diarize      Identify speaker segments in an audio file
  translate    Translate text
  synthesize   Synthesize speech to a WAV file
  bench        Benchmark an end-to-end ASR workload
  pull         Download a pinned model from Hugging Face
  model        List, pull, or inspect models
  doctor       Inspect runtime and device availability
  help         Show help for a command

Global options:
  -h, --help       Show help
  --version        Show version
  --json           Emit machine-readable results and errors
  --quiet          Suppress non-result progress messages
  --verbose        Emit additional diagnostics on stderr

使用できるモデルを見てみます。

$ build/metal-speech/bin/nemo-speech model list
Available models (* = command default)

ASR — transcribe, bench, serve
  * nemotron-3.5
      repo: nvidia/nemotron-3.5-asr-streaming-0.6b
      also: nemotron-asr
    nemotron-en
      repo: nvidia/nemotron-speech-streaming-en-0.6b
    parakeet-ctc
      repo: nvidia/parakeet-ctc-1.1b
    parakeet-tdt
      repo: nvidia/parakeet-tdt-0.6b-v3

Diarization — diarize, transcribe --diarize, serve
  * nemotron-3-diarization
      repo: nvidia/Nemotron-3-Diarization
      also: nemotron-diar
    sortformer
      repo: nvidia/diar_streaming_sortformer_4spk-v2
      also: sortformer-diar

TTS — synthesize, serve
  * magpie (speech model + tokenizer)
      repo: nvidia/magpie_tts_multilingual_357m
      also: magpie-tts
      pulls: nano-codec
  * nano-codec (codec)
      repo: nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
      also: nanocodec

Use a short name or full repository ID wherever MODEL is accepted.
Download ahead of time with: nemo-speech pull <name>

こちらではDiarizationではNemotron 3 Diarizationがデフォルトになっています。

$ build/metal-speech/bin/nemo-speech diarize /path/to/audio.wav
[nemo-speech] diarize session started
[model] downloading nvidia/Nemotron-3-Diarization@f667ed73aee5 (diarization, 102.1 MiB)
[model] license: NVIDIA Open Model License (OpenMDW 1.1) — https://huggingface.co/nvidia/Nemotron-3-Diarization
######################################################################################################################################################################################## 100.0%
[model] verifying size and SHA-256...
[model] ready: /Users/natsume.yuta/Library/Caches/NeMoSpeech/models/nvidia/Nemotron-3-Diarization/f667ed73aee57d40cc39428eb768b4fd87a0a29e/Nemotron-3-Diarization.q8_0.gguf
0.341   14.749  speaker 1
14.801  27.549  speaker 2
27.471  58.429  speaker 1

... (中略) ...

1209.321        1215.679        speaker 1
1215.441        1230.329        speaker 2
1230.201        1232.309        speaker 1
1232.691        1247.239        speaker 1
1247.931        1254.789        speaker 1
1255.101        1259.599        speaker 1
[nemo-speech] diarize session finished

Nemotron 3 Diarizationを使って話者識別ができました。

5. まとめ

以上、Nemotron 3 Diarizationを使ってみる話でした。

NeMo Speechで新しいリリースがあるまでは、モデルカードに記載された方法では使えなさそうです。
早いところ対応してほしいなぁと思います。

何かのお役に立てたら幸いです。

この記事をシェアする

関連記事