# Voice Pipeline This document describes the complete path of a voice frame from the transmitting device's microphone to the receiving device's speaker under the Floor Control Protocol (SCS-PTT-FLOOR-CONTROL-001). --- ## Transmit Side (pttclient — AudioTask) ``` PTT press │ ├─ Guard checks (any fail → abort, BUSY tone): │ ├─ socket.connected == true │ ├─ txStatus == false │ ├─ rxStatus == false │ ├─ networkStatus == CONNECTED (not SESSION_PENDING) │ └─ targetTxEnabled == true │ │ Increment talkspurtCounter (wraps at 65535) │ Generate talkspurtUUID (UUID v4) │ Sign talkspurtUUID → talkspurtSig (ML-DSA-65, device DSA private key) │ ├─ If private call (targetPeerDni set): │ ML-KEM-768 encapsulate(targetKemPublicKey) │ → kemCiphertext (1088 bytes) + 32-byte shared secret (AES-256 key) │ ML-DSA-65 sign(myDsaPrivKey, kemCiphertext) → kemSig │ key field = base64(kemCiphertext) + "|" + base64(kemSig) │ Store shared secret as pending TX AES-256-GCM session key │ ├─ If encrypted group call (group key available): │ key field = "GRP:v1:" │ ├─ If unencrypted group call: │ key field = omitted │ │ emit floor_request { │ v, tgid, srcDni, talkspurt, ts, │ talkspurtUUID, talkspurtSig, srcPubKey, key (if applicable) │ } with ack callback, 3-second timeout │ ├─ ack { ok: false } or timeout → ERROR tone, abort (no TX) │ └─ ack { ok: true }: │ Play TX start tone (TONE_CDMA_LOW_PBX_SSL if encrypted, TONE_CDMA_ALERT_INCALL_LITE if unencrypted; 100 ms) setTxStatus(true) Start AudioRecord │ ▼ Per frame loop: │ AudioRecord reads opusFrameBytes (320 bytes = 160 samples × 2) │ ▼ ByteBuffer → short[] (160 samples, little-endian) │ ▼ OpusEncoder.encode(samples, 0, 160) │ Returns variable-length Opus bitstream (typically 10–40 bytes) │ ▼ Encryption (one of three paths): Path A — Private call (AES session key established via floor_key_delivery): │ Per frame: │ IV: AESEncryption.getIVSecureRandomGCM("AES") (12 random bytes) │ AAD: "tg:" as UTF-8 │ Output: ciphertext || 16-byte GCM auth tag │ payload = base64(ciphertext + tag), iv = base64(IV) Path B — Group/broadcast call with group key (GRP:v1): │ secretKeySpec256 loaded from GroupKeyInventory │ Same AES-256-GCM per-frame as Path A │ AAD: "tg:" as UTF-8 │ key field = "GRP:v1:" included on every frame Path C — Unencrypted group call: │ payload = base64(VoiceHandler.encodeVoice(opusFrame, targetTGID, tgType)) │ iv omitted │ ▼ v2 envelope JSON (voice_data): { "v": 2, "tgid": , "srcDni": "", "seq": , ← starts at 0, no special meaning "talkspurt": , "ts": , "tgType": "G"|"B"|"P", "payload": "", "iv": "", ← omitted if unencrypted "key": "" ← group encrypted only, every frame } │ ▼ sendVoiceOrUdp(envelope): If isUdpPathAvailable() == true: UdpVoiceFrameBuilder.fromV2Envelope(...) → binary UDP frame (ustLen = 0, no sig/key/pubkey tags) → SKE wrap if skeActive() → UdpVoiceSocket.sendVoicePacket(packet) Else: socket.emit("voice_data", envelope) │ ▼ Repeat until PTT release or ptt_interrupt ``` **Talkspurt boundary:** `talkspurt` counter increments on every PTT press (wraps at 65535). Receivers use it alongside `srcDni` to detect the start of a new transmission and reset audio state. The `talkspurtUUID` is used only in `floor_request`; it does not appear in voice frames. **No preroll frame.** Voice capture begins immediately after the TX start tone. The first captured Opus frame is `seq == 0`. There is no silent preroll frame and no special seq 0 logic on either client or server. --- ## Floor Control — Key Delivery Path (Private Calls) For private calls, the AES session key reaches the receiver before any audio frame. This sequence completes inside the `floor_request` exchange: ``` pttclient (sender) ptt-server pttclient (target) │ │ │ │ floor_request │ │ │ (key = ML-KEM-768 ct │ │ │ + ML-DSA-65 sig) │ │ │ ────────────────────────► │ │ │ │ floor_key_delivery │ │ │ { senderDni, uuid, key } │ │ │ ─────────────────────────►│ │ │ │ ML-KEM decapsulate │ │ │ Store AES-256 key by senderDni │ │ ack { ok: true } │ │ │ ◄─────────────────────────│ │ ack { ok: true } │ │ │ ◄─────────────────────────│ │ │ │ │ Play TX start tone │ │ Start capture │ │ First voice frame (seq 0) ──► │ ──────────────────────────►│ Decrypt immediately ``` The receiver holds the session key before the first audio frame arrives. There is no dependency on seq 0 for key material. A narrow pending queue exists solely for UDP packet reordering (`drainPendingPrivateVoiceAfterSessionKey`). --- ## Server Routing (ptt-server) ``` floor_request received from socket │ │ 1. Verify srcDni == socket.data.dni │ 2. Derive expected DNI from srcPubKey: │ 1952-byte raw key → deriveDniFromRawBytes() (ML-DSA-65) │ → SIG_FAIL if derived DNI != srcDni │ 3. verifyTalkspurtSig(srcPubKey, talkspurtUUID, talkspurtSig) │ ML-DSA-65 → mlDsa65Verify() │ → SIG_FAIL if invalid │ 4. Rate limit check │ ├─ If tgid is DNI (private call): │ Resolve target socket from dev: room │ Emit floor_key_delivery to target (3 s timeout) │ Await target ack │ Return ack to sender │ └─ If tgid is integer (group/broadcast): Check broadcastAuth if type == 'B' Store talkspurt context { srcDni, talkspurtUUID, tgid, ts } Return ack { ok: true } to sender voice_data received from socket (post floor_request): │ │ 1. Verify srcDni == socket.data.dni → set env.src │ 2. Check seq vs PTT_MAX_VOICE_SEQ │ If exceeded and tgType != 'B' → emit ptt_interrupt, discard │ 3. Check mayTransmitGroup (group/broadcast) │ 4. Route: │ DNI target → socket.to("dev:").emit('voice_data', env) │ Group target → socket.to("grp:").emit('voice_data', env) │ socket.to("scan-").emit('voice_data', env) │ │ NO: sig verification, pubkey derivation, UST check, seq 0 gate, │ session key extraction, talkspurtMeta lookup ``` --- ## Receive Side (pttclient — SocketService) ### TCP receive path ``` socket.on("voice_data", envelope) │ │ 1. Parse tgid, srcDni, seq, talkspurt, tgType │ 2. Route check: is this for me? (my group, my DNI, scan list) │ → discard if not routed to this device │ 3. Priority and scan filtering │ → discard lower-priority concurrent transmissions │ 4. Decode payload: │ base64Decode("payload") → ciphertext+tag or raw audio │ base64Decode("iv") → 12-byte GCM nonce (if present) │ ├─ Encrypted private frame (iv present, no key field): │ Look up session key by srcDni from rxPrivateKeyMap │ (populated by floor_key_delivery handler) │ AAD = "tg:" as UTF-8 │ AES-256-GCM decrypt → raw Opus bitstream │ ├─ Encrypted group frame (iv present, key = "GRP:v1:"): │ Look up AES-256 key from GroupKeyInventory by kid │ AAD = "tg:" as UTF-8 │ AES-256-GCM decrypt → raw Opus bitstream │ └─ Unencrypted frame (iv absent): payload is raw encoded audio → VoiceHandler.decodeVoice(...) → raw Opus bitstream │ ▼ OpusDecoder.decodeToByteArray(opusBitstream) → PCM byte[] │ ▼ playbackQueue.offer(PlaybackFrame) │ ▼ Playback thread: AudioTrack.play() + write(pcmBytes) ``` ### UDP receive path ``` UdpVoiceSocket receive loop: │ Receive datagram from channel │ ├─ If SkeOuterEnvelope.looksLikeSke() and skeActive(): │ SkeOuterEnvelope.tryUnwrap(datagram, mSkeKey) → inner bytes │ └─ Parse inner frame type: ├─ TYPE_PROBE_ACK (0x05): set mUdpReachable, notify Listener.onUdpReachable() └─ TYPE_RELAY (0x03): UdpVoiceFrameParser.parse(inner) → UdpVoiceFrame │ handleUdpVoiceFrame(frame): Reconstruct v2 JSON envelope from UdpVoiceFrame fields Call handleVoiceDataV2(env, "UDP") → same decrypt/decode/playback path as TCP receive ``` The UDP receive path is identical in behavior to the TCP receive path from the decryption step onward. There is no seq 0 special case in either path. --- ## floor_key_delivery handler (pttclient receive side) ``` socket.on("floor_key_delivery", payload) │ │ Extract senderDni, talkspurtUUID, key, ts │ Validate senderDni is a valid 16-digit DNI │ Check talkspurtUUID not previously received from senderDni (replay) │ → ack { ok: false } if replay │ │ Split key on "|" → kemCtB64, sigB64 │ base64Decode(kemCtB64) → kemCiphertext (1088 bytes) │ Verify ML-DSA-65: mlDsaVerify(senderDsaPubKey, kemCiphertext, sigBytes) │ → reject if invalid and sender DSA key is known │ ML-KEM-768 decapsulate(myKemPrivKey, kemCiphertext) → 32-byte shared secret │ Use shared secret directly as AES-256-GCM session key │ Store in rxPrivateKeyMap[senderDni] (replace any prior entry) │ Record talkspurtUUID in seen-UUID set for senderDni │ └─ ack { ok: true } ``` --- ## Audio Specifications | Parameter | Value | |---|---| | Sample rate | 8000 Hz | | Channels | Mono (1) | | Bit depth | 16-bit signed PCM | | Frame size | 160 samples = 20 ms | | Codec | Opus (libopus via JNI) | | Encryption | AES-256-GCM, 12-byte IV, 128-bit auth tag | | AAD | `"tg:"` as UTF-8 (decimal group ID or 16-digit peer DNI) | | Session key lifetime (private) | One PTT press (new key generated per talkspurt, delivered via floor_key_delivery) | | Session key lifetime (group) | Per GroupKeyInventory entry validity period | | Max frames per PTT | `PTT_MAX_VOICE_SEQ` (default 3000 ≈ 60 seconds at 20 ms/frame) | --- ## Removed Mechanics (Deprecated) The following mechanisms from the v1.x pipeline are removed and MUST NOT be implemented: | Removed mechanism | Was located in | Replacement | |---|---|---| | seq 0 silent preroll frame | `AudioTask` preroll block | None — floor_request provides key delivery | | `savedPrerollTalkspurtUuid` state | `AudioTask` / `SocketService` | `talkspurtUUID` lives in floor_request context | | `pendingPrivateVoiceUntilKeyTalkspurt` queue (bulk pre-key buffering) | `SocketService.handleVoiceDataV2` | Narrowed to UDP reorder only (`drainPendingPrivateVoiceAfterSessionKey`); no longer used for holding frames until seq 0 | | `handleSessionKeyV2` (seq 0 key extraction) | `SocketService.handleVoiceDataV2` | `floor_key_delivery` listener | | seq 0 talkspurtSig / srcPubKey verify in voice handler | `SocketService.handleVoiceDataV2` | Performed in server `floor_request` handler | | seq 0 UST validation in UDP relay | `udp-voice-relay.js` | PROBE / PROBE_ACK is sole binding mechanism | | `talkspurtMeta` map in UDP relay | `udp-voice-relay.js` | Per-frame tgType from extension block | | Extension tags 0x01–0x03 in UDP frames | `UdpVoiceFrameBuilder` / `UdpVoiceFrameParser` | Removed | | UST bytes in voice frames | `UdpVoiceFrameBuilder` | `ustLen` always 0; UST in PROBE only |