promotional bannermobile promotional banner

Speaker API

TEXT→HUMAN VOICE in Minecraft.

Speaker API is an offline, privacy-first text-to-speech (TTS) middleware for Minecraft. It converts in-game text — chat, system messages, toasts, and any text submitted by other mods — into natural, human-like speech played through your speakers. Everything runs locally on your machine: no network requests, no cloud, no accounts.

What it does

  • Local, offline TTS powered by sherpa-onnx and Kokoro voice models (both Apache-2.0).
  • Chinese-first reading with mixed Chinese/English support and multiple built-in voices.
  • Other mods can register their own "sources" and submit text through a simple API to receive spoken feedback.
  • Optional: replace the system narrator (Windows Narrator / SAPI) so other mods and the game itself speak through this mod instead.
  • Position-aware audio: bind speech to an entity or block so the voice appears to come from that location in the world.

Why you might want it

  • Fully offline and private — your text never leaves your computer.
  • Gives Minecraft, and accessibility-focused mods, a real, natural-sounding voice instead of silent text.
  • Lightweight and client-side only; it does not affect game servers.

Important before you download

  • Client-side only. Install on the Minecraft client (Forge / NeoForge), not on a server.
  • On first use it downloads a voice model (about 380 MB). A download-mirror option is included for faster fetching; you can also run /speakerapi download kokoro-multilang.
  • Supported versions: Minecraft 1.19.2 (Forge), 1.20.1 (Forge), 1.21.1 (NeoForge).
  • Requires Java 17 for the Forge builds, Java 21 for the NeoForge build.

Commands

  • /speakerapi test — speak a sample sentence (also loads the model if needed).
  • /speakerapi status — show engine and model status in chat.
  • /speakerapi model <id> — select and load a model.
  • /speakerapi download <id> — download a model only.
  • /speakerapi stop — stop all speech.
  • /speakerapi voices — list available voices.

Credits

  • Speech engine: sherpa-onnx (Apache-2.0).
  • Voice models: Kokoro (Apache-2.0) and other open-source models.
  • Author: li2012China.

API Usage (for other mod developers)

Speaker API exposes a stable, cross-version API under the package net.speakerapi. All methods are static. The mod is client-side only: API calls made from a dedicated/server environment are safe no-ops (they simply do nothing), so guard with isEnabled() / isEngineReady() if you also run headless.

Adding the dependency

On Forge / NeoForge, add a dependency on speakerapi (any supported version). At runtime the class net.speakerapi.SpeakerAPI is available whenever the mod is loaded.

Basic speech

import net.speakerapi.SpeakerAPI;
import net.speakerapi.SpeakerAPI.Priority;
import net.speakerapi.core.VoiceProfile;
import net.minecraft.network.chat.Component;

// Speak a line from your mod ("mymod" is the source id).
SpeakerAPI.say(Component.literal("Hello, world"), "mymod", VoiceProfile.ZH_FEMALE_NATURAL, Priority.NORMAL);

// Or with a plain string:
SpeakerAPI.say("Watch your step", "mymod", VoiceProfile.ZH_MALE_NATURAL, Priority.HIGH);
  • sourceId — string identifying your mod; used for per-source filtering/registration.
  • voice — a VoiceProfile. Built-ins: VoiceProfile.ZH_FEMALE_NATURAL, VoiceProfile.ZH_MALE_NATURAL. Discover more via availableVoices() / voiceFor("zh_female_natural").
  • priority — Priority.LOW | NORMAL | HIGH | ANNOUNCE.

Cancellation & completion

The say(Component/String, …) overloads return CompletableFuture<Void> (completes when speech finishes). The positioned / builder variants return a SpeakHandle:

SpeakerAPI.SpeakHandle handle = SpeakerAPI.sayForEntity(player, "Follow me");
handle.cancel();                              // stop this utterance early
boolean done = handle.isDone();
handle.future().thenRun(() -> System.out.println("finished"));

Positioned & bound (follow) audio

// One-shot spatialized at a coordinate / entity / block (snapshot at call time):
SpeakerAPI.sayAt("Danger zone", "mymod", null, Priority.HIGH, someBlockPos);
SpeakerAPI.sayAt("Enemy approaching", "mymod", null, Priority.HIGH, someEntity);

// Bound audio: the voice follows a moving entity for the whole utterance:
SpeakerAPI.SpeakHandle h = SpeakerAPI.sayForEntity(player, "Come with me");
SpeakerAPI.SpeakHandle b = SpeakerAPI.sayForBlock(blockPos, "This is the workbench");

For full control, use the builder:

SpeakerAPI.SayOptions opt = SpeakerAPI.SayOptions.builder()
    .text("Quest complete")
    .sourceId("mymod")
    .voice(VoiceProfile.ZH_FEMALE_NATURAL)
    .priority(Priority.ANNOUNCE)
    .volume(0.8f)                                       // optional per-call gain
    .category(net.minecraft.sounds.SoundSource.RECORDS) // optional MC sound category
    .bindTo(targetEntity)                               // follow the entity (eye height)
    .onSentenceStart(s -> { /* per-sentence */ })
    .onSentenceDone(s -> { /* per-sentence */ })
    .onComplete(() -> { /* whole utterance done */ })
    .build();
SpeakerAPI.say(opt);

bindTo(Entity) / bindTo(BlockPos) make the voice track the target; you can also pass a custom PositionProvider (net.speakerapi.core.PositionProvider) for arbitrary moving anchors. Spatialization needs a client-side listener pose; if unavailable it degrades to centered playback (still audible).

Engine readiness

The engine loads the voice model asynchronously on first use. Check or wait before relying on speech:

if (SpeakerAPI.isEngineReady()) {
    // speak now
}
SpeakerAPI.addEngineReadyListener(() -> { /* engine loaded */ });

Models (select / download)

SpeakerAPI.knownModelIds();                 // supported model ids
SpeakerAPI.selectModel("kokoro-multilang"); // set default + download if missing + hot-reload
SpeakerAPI.downloadModel("kokoro-multilang"); // download only (CompletableFuture<Boolean>)
SpeakerAPI.listModels();                    // local readiness snapshot

Other controls

SpeakerAPI.setEnabled(boolean);          // global mute toggle
SpeakerAPI.stopAll();                    // stop every active utterance
SpeakerAPI.setCategory(SoundSource);     // bind output to a MC sound category
SpeakerAPI.setReplaceSystemTts(boolean); // route the system narrator through this mod
SpeakerAPI.registerSource("mymod", text -> true); // register a named source

Notes

  • Client-side only. All speech is local; no text leaves the machine.
  • The API surface is identical across Minecraft 1.19.2 / 1.20.1 (Forge) and 1.21.1 (NeoForge).

Speaker API Team

profile avatar
  • 2
    Projects
  • 12
    Downloads

More from li2012China