Skip to content
Merged
Show file tree
Hide file tree
Changes from 8 commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
7c891f6
docs: plan live dictation and Soniox bypass
Snehit70 Jun 26, 2026
7d7ca92
feat: add live dictation text writer
Snehit70 Jun 26, 2026
b7c41de
feat: add live dictation config
Snehit70 Jun 26, 2026
42adde3
feat: add live provider transcript contract
Snehit70 Jun 26, 2026
55e684f
feat: wire live dictation transcript events
Snehit70 Jun 26, 2026
6bf20c6
feat: add soniox streaming provider
Snehit70 Jun 26, 2026
d33d824
feat: wire soniox live bypass hotkey
Snehit70 Jun 26, 2026
0dc530b
docs: document live dictation configuration
Snehit70 Jun 26, 2026
57bed1e
feat(config): add retypeFormatted and paragraphPauseMs to live dictat…
Snehit70 Jun 26, 2026
b9425c6
feat(soniox): add token spacing, paragraph breaks, LLM formatting, an…
Snehit70 Jun 26, 2026
937b218
feat(service): wire LLM formatting, retype mechanism, and IPC soniox-…
Snehit70 Jun 26, 2026
1ad204f
feat(cli): add soniox-toggle subcommand via IPC socket
Snehit70 Jun 26, 2026
ad0a16b
test(soniox): add paragraph break, spacing, and tag-stripping tests
Snehit70 Jun 26, 2026
4541f60
docs: document Soniox formatting pipeline, config options, and PRD
Snehit70 Jun 26, 2026
d04966d
fix(soniox): use join('') for tokens that already include spaces
Snehit70 Jun 26, 2026
eb70101
fix(soniox): strip spaces before punctuation in live tokens
Snehit70 Jun 26, 2026
53cfab4
fix(live-dictation): use Shift+End escape sequence on Wayland
Snehit70 Jun 26, 2026
a48050b
fix(live-dictation): use wtype native key simulation for select on Wa…
Snehit70 Jun 26, 2026
61a02f5
feat(soniox): pass boost words to Soniox as context.terms
Snehit70 Jun 26, 2026
02ffa1d
feat(soniox): add context.general, contextText, languageHintsStrict c…
Snehit70 Jun 27, 2026
fba140b
feat(soniox): send context.general, contextText, languageHintsStrict …
Snehit70 Jun 27, 2026
fdaea01
test: add tests for Soniox context config fields; update docs
Snehit70 Jun 27, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,14 @@ _Avoid_: hotkey, shortcut
The on-screen status surface that reflects live daemon state.
_Avoid_: popup, HUD

**Live Dictation**:
Entering stable transcript text into the currently focused text field while a recording is still active, then preserving the final transcript through the normal clipboard and history paths.
_Avoid_: live paste, streaming paste

**Provider Bypass**:
A recording path that uses one live transcription provider directly and skips the Groq plus Deepgram merge and quality pipeline by design.
_Avoid_: fallback, fast mode

## Observability

**Readiness**:
Expand Down
39 changes: 38 additions & 1 deletion docs/CONFIGURATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,7 @@ If an API key is missing from `config.json`, the application will fall back to t
- `GROQ_API_KEY`: Fallback for `apiKeys.groq`
- `GROQ_FALLBACK_API_KEY`: Fallback for `apiKeys.groqFallback`
- `DEEPGRAM_API_KEY`: Fallback for `apiKeys.deepgram`
- `SONIOX_API_KEY`: Fallback for `apiKeys.soniox` when Soniox live dictation is enabled

### Logging
- `LOG_LEVEL`: Sets the minimum logging level. Options: `trace`, `debug`, `info`, `warn`, `error`, `fatal`, `silent`. Default: `info`.
Expand All @@ -44,7 +45,8 @@ The configuration is a JSON file structured into several sections.
"apiKeys": {
"groq": "gsk_...",
"groqFallback": "gsk_...",
"deepgram": "00000000-0000-0000-0000-000000000000"
"deepgram": "00000000-0000-0000-0000-000000000000",
"soniox": "soniox_..."
},
"behavior": {
"hotkey": "Right Control",
Expand Down Expand Up @@ -87,6 +89,14 @@ The configuration is a JSON file structured into several sections.
"Groq",
"Deepgram"
]
},
"liveDictation": {
"enabled": false,
"insertionCommand": "auto",
"soniox": {
"enabled": false,
"triggerKey": "Right Alt"
}
}
}
```
Expand All @@ -104,6 +114,7 @@ Authentication credentials for the transcription services.
| `groq` | String | N/A | API key for Groq (Whisper V3). | Must start with `gsk_`. | [Groq Console](https://console.groq.com/keys) |
| `groqFallback` | String | Optional | Secondary Groq API key used only when the primary merge key is rate-limited/quota-limited. | Must start with `gsk_` if provided. | [Groq Console](https://console.groq.com/keys) |
| `deepgram` | String | N/A | API key for Deepgram (Nova-3). | 40-char hex string or UUID. | [Deepgram Console](https://console.deepgram.com/) |
| `soniox` | String | Optional | API key for Soniox real-time STT. Required only when `liveDictation.soniox.enabled` is used. | Non-empty string. | [Soniox Console](https://console.soniox.com/) |

#### How to obtain API Keys

Expand All @@ -117,6 +128,11 @@ Authentication credentials for the transcription services.
- Navigate to **API Keys** and create a new key.
- **Format**: The key is typically a **40-character hexadecimal string** (e.g., `abcdef1234567890abcdef1234567890abcdef12`). Legacy keys or specific project IDs might use a UUID format, both are supported.

3. **Soniox API Key**:
- Go to the [Soniox Console](https://console.soniox.com/).
- Create an API key for real-time speech-to-text.
- Hyprvox reads it from `apiKeys.soniox` or `SONIOX_API_KEY`.

---

### 2. Behavior (`behavior`)
Expand Down Expand Up @@ -300,6 +316,27 @@ In streaming mode, performance logs include Deepgram finalization observability:

These fields are for analysis and future tuning. Do not treat a frequent `finalize_timeout` by itself as proof that endpointing should change; compare the final transcript quality and late finalization signals first.

### 5. Live Dictation (`liveDictation`)

Live Dictation controls focused-text insertion while recording. It is disabled by default.

| Option | Type | Default | Description | Validation Rules |
| :--- | :--- | :--- | :--- | :--- |
| `enabled` | Boolean | `false` | Type stable live transcript text into the currently focused input during the normal streaming path. | Requires `transcription.streaming: true` to receive Deepgram live transcript events. |
| `insertionCommand` | String | `"auto"` | Command used for focused text insertion. | `"auto"`, `"wtype"`, or `"xdotool"`. |
| `soniox.enabled` | Boolean | `false` | Enable the separate Soniox provider-bypass hotkey. | Requires `apiKeys.soniox` or `SONIOX_API_KEY` when used. |
| `soniox.triggerKey` | String | `"Right Alt"` | Separate hotkey for Soniox live dictation. | Same hotkey format as `behavior.hotkey`. |

#### Normal Live Dictation

When `liveDictation.enabled` and `transcription.streaming` are both enabled, Hyprvox types committed Deepgram streaming transcript chunks into the focused input as they arrive. The final transcript still follows the normal Groq plus Deepgram merge, clipboard, history, and validation path after recording stops.

#### Soniox Provider Bypass

When `liveDictation.soniox.enabled` is enabled, pressing `liveDictation.soniox.triggerKey` starts a Soniox real-time STT session. Recorder PCM is sent directly to Soniox, stable final tokens are typed into the focused input, and the final Soniox transcript is copied to clipboard and appended to history when recording stops.

This path intentionally skips Groq, Deepgram batch transcription, merge, repair, and quality recovery. Use it when low-latency live dictation is more important than the normal multi-provider quality pipeline.

#### Language Options

For **v1.0**, `hyprvox` is optimized for and officially supports **English only**.
Expand Down
101 changes: 101 additions & 0 deletions docs/ISSUES-LIVE-DICTATION-SONIOX.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
# Offline Issue Breakdown: Live Dictation And Soniox Provider Bypass

## 1. Live Dictation Text Writer Contract

**Type**: AFK

**Blocked by**: None - can start immediately

**User stories covered**: 1, 3, 4, 5, 6, 7, 8, 15, 18

## What to build

Build the smallest focused-input writer path that can accept stable transcript events and type only the committed delta into the currently focused field. The slice should include a command-backed text injection boundary with test doubles for command execution.

## Acceptance criteria

- [ ] Stable transcript chunks insert only new committed text.
- [ ] Repeated transcript events do not duplicate already inserted text.
- [ ] Wayland uses `wtype` by default when available.
- [ ] X11 uses `xdotool type` when Wayland is not active.
- [ ] Tests do not type into the real desktop.
- [ ] Insertion errors are returned as structured failures that callers can log and recover from.

## 2. Live Dictation Config And Readiness

**Type**: AFK

**Blocked by**: Issue 1

**User stories covered**: 8, 9, 13, 14, 18

## What to build

Add configuration for Live Dictation and Soniox credentials without changing default transcription behavior. Readiness checks should report missing optional dependencies only when the feature is enabled.

## Acceptance criteria

- [ ] Existing minimal configs still parse and keep Live Dictation disabled by default.
- [ ] Live Dictation config validates trigger key and insertion command settings.
- [ ] Soniox credentials can be read from config or environment.
- [ ] Missing Soniox credentials do not break default Groq plus Deepgram recording.
- [ ] Tests cover config defaults, enabled Live Dictation config, and invalid trigger key handling.

## 3. Live Provider Contract Around Existing Deepgram Streaming

**Type**: AFK

**Blocked by**: Issue 1

**User stories covered**: 1, 2, 9, 16, 17, 18

## What to build

Introduce a live provider contract that represents streaming transcript events, PCM input, and final transcript stop behavior. Adapt the current Deepgram streaming path to that contract while preserving current default behavior.

## Acceptance criteria

- [ ] Normal streaming recordings still produce the same final clipboard behavior.
- [ ] Live provider events can feed the Live Dictation text writer when enabled.
- [ ] Deepgram streaming failures still fall back to existing batch behavior.
- [ ] Tests prove the contract with a fake provider and do not call Deepgram.

## 4. Soniox Provider Bypass Recording Path

**Type**: AFK

**Blocked by**: Issues 1, 2, 3

**User stories covered**: 10, 11, 12, 13, 14, 16, 17, 18

## What to build

Add Soniox as a live provider and wire a separate provider-bypass trigger key. The path streams recorder PCM to Soniox, feeds stable transcript events to Live Dictation, and writes the final Soniox transcript to clipboard and history without running Groq, Deepgram batch, merge, repair, or quality recovery.

## Acceptance criteria

- [ ] Soniox provider connects to the documented real-time STT WebSocket endpoint.
- [ ] Provider bypass does not call Groq, Deepgram batch, merge, repair, or quality recovery.
- [ ] Final Soniox transcript is copied to clipboard and appended to history.
- [ ] Provider errors surface as user-facing Soniox errors.
- [ ] Tests use a fake WebSocket transport and do not hit Soniox.

## 5. Runtime Verification And PR Readiness

**Type**: HITL

**Blocked by**: Issues 1, 2, 3, 4

**User stories covered**: 1-18

## What to build

Run the focused local test suite, verify no default behavior changed, perform a manual live dictation check on the user machine, then prepare incremental commits and a PR.

## Acceptance criteria

- [ ] Focused unit tests pass.
- [ ] Default config and normal transcription behavior remain unchanged.
- [ ] Live Dictation can be manually tested in a scratch focused input.
- [ ] Soniox provider bypass can be manually tested when credentials are available.
- [ ] Commits are incremental and scoped to planning, infrastructure, provider contract, Soniox path, and verification.
77 changes: 77 additions & 0 deletions docs/PRD-LIVE-DICTATION-SONIOX.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
# Live Dictation And Soniox Provider Bypass PRD

## Problem Statement

Hyprvox currently optimizes for final transcript quality: the user records, waits for provider results, receives a merged transcript, and then pastes from the clipboard. That is accurate, but it is not enough for workflows where the focused input should fill as the user speaks. The user also wants a low-latency Soniox live path that bypasses the existing Groq plus Deepgram merge pipeline while still copying the final transcript to the clipboard after the recording ends.

## Solution

Add Live Dictation as an explicit output mode. During an active recording, stable live transcript text is typed into the currently focused text field. When the recording stops, Hyprvox still writes the final transcript to clipboard and history.

Add Soniox as a live transcription provider. Soniox provider bypass mode starts a recording through a separate trigger key, streams PCM to Soniox, types stable live transcript text into the focused text field, and copies the final Soniox transcript to the clipboard. It intentionally skips Groq, Deepgram batch transcription, merge, repair, and quality recovery.

## User Stories

1. As a Hyprvox user, I want dictated words to appear in the focused input while I speak, so that I do not have to wait until recording stops to see text.
2. As a Hyprvox user, I want the final transcript copied to the clipboard after live dictation ends, so that my existing paste workflow still works.
3. As a Hyprvox user, I want live text insertion to avoid duplicate partial phrases, so that the focused input remains readable.
4. As a Hyprvox user, I want live insertion to type only stable transcript text, so that unstable interim provider guesses do not corrupt the input.
5. As a Hyprvox user, I want live insertion to use the current focused text field, so that the feature works across editors, browsers, terminals, and chat inputs.
6. As a Hyprvox user on Wayland, I want live insertion to use the available desktop text injection tool, so that the feature works in my normal Hyprland session.
7. As a Hyprvox user on X11, I want a compatible text injection fallback, so that live dictation is not Wayland-only.
8. As a Hyprvox user, I want live insertion failures to fall back to normal clipboard behavior, so that transcription is not lost.
9. As a Hyprvox user, I want normal Groq plus Deepgram transcription to keep working unchanged unless live dictation is enabled, so that quality-critical recordings are not affected.
10. As a Hyprvox user, I want a separate Soniox trigger key, so that I can choose a low-latency live provider without changing the default trigger key behavior.
11. As a Hyprvox user, I want Soniox provider bypass to skip merge and quality recovery, so that the final transcript reflects the live provider directly with low latency.
12. As a Hyprvox user, I want Soniox provider bypass to still write the final transcript to clipboard and history, so that downstream workflows remain consistent.
13. As a Hyprvox user, I want Soniox authentication to be configured through Hyprvox config or environment variables, so that it behaves like the existing provider keys.
14. As a Hyprvox user, I want Soniox provider errors to be reported clearly, so that I can distinguish auth, network, and runtime failures.
15. As a Hyprvox maintainer, I want the live text writer to be tested independently from real desktop input, so that the behavior is stable without disrupting laptop usage.
16. As a Hyprvox maintainer, I want provider bypass tests to exercise observable daemon behavior, so that refactors do not break the user-facing modes.
17. As a Hyprvox maintainer, I want live provider implementations behind a small interface, so that Deepgram and Soniox streaming can share daemon orchestration without coupling to provider-specific SDK details.
18. As a Hyprvox maintainer, I want benchmark thresholds for live insertion and stop latency, so that the feature has measurable quality gates.

## Implementation Decisions

- Keep normal recording semantics intact. The default trigger key continues to run the existing Groq plus Deepgram path unless the user explicitly enables Live Dictation for that path.
- Introduce a Live Dictation output component that accepts transcript events and emits text insertion operations. It tracks committed text and only inserts stable deltas.
- Treat provider interim text as display/input candidates and provider final text as committed text. Stable final chunks are the only text typed by default.
- Use `wtype` as the primary Wayland focused-input writer because it is installed on the target machine and matches the Hyprland environment. Use `xdotool type` as an X11 fallback. Keep `ydotool` as a later fallback because it may require daemon permissions.
- Add configuration for Live Dictation enablement, insertion command selection, and Soniox provider credentials.
- Add a separate Soniox provider bypass trigger key. This mode has its own lifecycle but reuses audio capture, PCM streaming, clipboard output, history, notifications, and daemon state where practical.
- Add a small live provider contract with start, send PCM, stop, and transcript events. Deepgram streaming can be adapted to that contract, and Soniox can implement the same contract.
- Soniox provider bypass uses Soniox real-time STT over WebSocket. Current official docs describe real-time STT as WebSocket-based, with endpoint `wss://stt-rt.soniox.com`; the SDK reference exposes `wss://stt-rt.soniox.com/transcribe-websocket` as the default STT WebSocket URL.
- Provider bypass final text is not sent through the Groq plus Deepgram merge result pipeline. It is copied as the Soniox final transcript and marked as provider bypass in history/metrics.
- Keep GitHub issue publishing out of scope for the first planning artifact. Issues are drafted offline in this repository.

## Testing Decisions

- Tests should verify behavior through public interfaces and shell/system boundaries, not private methods.
- Start with the Live Dictation text accumulator and text writer contract because it is the riskiest user-visible behavior and can be tested without real desktop input.
- Mock only external boundaries: text injection command execution, WebSocket transport, and provider API responses.
- Add focused tests for config parsing so Soniox and Live Dictation settings are validated without starting the daemon.
- Add provider-bypass tests around a fake live provider to prove final transcript copy behavior without hitting Soniox.
- Avoid local tests that open real microphones, real Electron windows, or real desktop notifications by default. Use existing `HYPRVOX_TEST_MODE=1` behavior and targeted unit tests first.

## Benchmarks And Quality Gates

- Live insertion latency: committed provider chunks should be scheduled for focused-input insertion within 100 ms of receipt in unit-level timing tests.
- Duplication guard: repeated partial/final events must not insert duplicate words.
- Stop-to-clipboard latency for provider bypass: no merge or batch provider call may run in the Soniox bypass path.
- Normal path safety: default Groq plus Deepgram recording tests must not observe any Soniox dependency or live insertion side effect when the feature is disabled.
- Desktop safety: automated tests must not type into the real desktop, send real notifications, or require real microphone input.
- Config safety: missing Soniox credentials must fail only when Soniox bypass mode is used or explicitly verified, not when normal transcription runs.

## Out of Scope

- Replacing the Groq plus Deepgram merge result pipeline.
- Streaming unstable interim text with later destructive edits in arbitrary inputs.
- Overlay redesign.
- Publishing PRD or issue drafts to GitHub before user approval.
- Tuning endpointing or merge models for existing providers.
- Supporting every desktop environment’s preferred text injection tool in the first slice.

## Further Notes

- Official Soniox references used for planning: `https://soniox.com/docs/api-reference/stt/websocket-api`, `https://soniox.com/docs/api-reference`, and `https://soniox.com/docs/stt/models`.
- The term `Provider Bypass` is intentionally distinct from `Fallback`: fallback is an error recovery behavior, while provider bypass is a deliberate low-latency mode.
9 changes: 9 additions & 0 deletions docs/STT_FLOW.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,7 @@ sequenceDiagram
- **Key**: Default is `Right Control`.
- **Behavior**: Toggle mode. The first press starts the recording; the second press stops it.
- **Verification**: Checks for hotkey conflicts at startup.
- **Soniox bypass**: If `liveDictation.soniox.enabled` is true, `liveDictation.soniox.triggerKey` starts a separate Soniox live dictation path instead of the normal Groq plus Deepgram path.

### 2. Audio Capture
- **Utility**: `arecord` (via `node-record-lpcm16`).
Expand Down Expand Up @@ -94,6 +95,14 @@ Before provider calls, the daemon builds technical-term hints from configured bo

**Live Groq Chunk Metrics**: Live chunking perf events include chunk counts, finalize wait, final-tail length, background request time, dropped/recovered chunk counts, and quality fallback status. `groqLiveDroppedChunks` tracks chunk-local outputs removed before stitching; `groqLiveRecoveredChunks` tracks dropped chunks recovered by contextual repair; `groqLiveQualityFallback` records when a clean-looking live Groq transcript was materially shorter than Deepgram and Hyprvox ran one full-audio Groq request before merge.

### 4b. Live Dictation And Soniox Provider Bypass

Live Dictation is an opt-in output path for committed streaming transcript chunks:

- In the normal path, `liveDictation.enabled` plus `transcription.streaming` types committed Deepgram transcript chunks into the focused input while recording continues. Stop still runs the normal Groq plus Deepgram merge, validation, clipboard, and history flow.
- In Soniox provider bypass mode, `liveDictation.soniox.triggerKey` starts a Soniox real-time WebSocket session. Recorder PCM is sent directly as 16 kHz mono `pcm_s16le`, final Soniox tokens are typed into the focused input, and the final Soniox transcript is copied to clipboard and history when recording stops.
- Soniox provider bypass intentionally skips Groq, Deepgram batch transcription, merge, repair, and quality recovery. The history engine is recorded as `soniox`, and aggregate performance logs use `mergeStrategy: "provider_bypass"`.

### 5. Merge And Quality Pipeline
If both Groq and Deepgram return results, Hyprvox first decides whether deterministic selection is enough or whether an LLM merge is needed. The merge model is configured by `transcription.mergeModel`.

Expand Down
Loading