feat: v0.1.0 — client Sailfish per hermes-webui con voce
- ApiClient HTTP/SSE: login, sessioni, chat/start, stream SSE, transcribe, tts - ChatModel: streaming live + risincronizzazione dallo stato del server - CookieJar persistente (0600, mai la password); Settings - Recorder OGG/Opus (fallback WAV/PCM, retry); Player TTS (QMediaPlayer) - UI Silica: splash, login, sessioni, chat (dettatura + read-aloud), settings, cover - Sailjail Internet;Audio;Microphone; icone; spec RPM; docs PROTOCOL/BUILD - tests/core_test: unit parser SSE + integrazione live (probe :8899)
This commit is contained in:
@@ -0,0 +1,93 @@
|
||||
# Build: compilazione, test e deploy
|
||||
|
||||
## 1. Prerequisiti SDK
|
||||
|
||||
- Sailfish SDK (testato su 5.1.0.11) con target del dispositivo aarch64:
|
||||
`SailfishOS-5.0.0.62-aarch64` (Jolla Phone 2026) o equivalente.
|
||||
- Nel target SDK serve **QtMultimedia** (registrazione/riproduzione):
|
||||
se `pkgconfig(Qt5Multimedia)` non è trovato in fase di build, installa il
|
||||
pacchetto devel nel target:
|
||||
|
||||
```bash
|
||||
sfdk tools exec SailfishOS-5.0.0.62-aarch64 zypper install -y qt5-qtmultimedia-devel
|
||||
```
|
||||
|
||||
(oppure via Qt Creator: Tools → Sailfish OS → Manage targets → Packages,
|
||||
oppure `sfdk target install` con un RPM locale.)
|
||||
|
||||
## 2. Build e deploy dell'app
|
||||
|
||||
```bash
|
||||
cd harbour-hermes
|
||||
sfdk build # genera RPMS/
|
||||
sfdk deploy # installa sul device collegato
|
||||
```
|
||||
|
||||
Note:
|
||||
- Il target è Qt 5.6: il codice del core è scritto per compilare anche lì
|
||||
(niente API Qt > 5.6; per il TTS si usa `audio/mpeg` di edge-tts).
|
||||
- Il profilo SailJail è nel `rpm/harbour-hermes.desktop`
|
||||
(`Internet;Audio;Microphone`) e deve restare allineato a
|
||||
`setOrganizationName("harbour")` / `setApplicationName("hermes")` in `main.cpp`.
|
||||
- Al primo uso del microfono Sailfish chiede il consenso al permesso
|
||||
**Microphone**.
|
||||
|
||||
## 3. Test del core in locale (senza SDK)
|
||||
|
||||
Il core (ApiClient, ChatModel, parser SSE, Recorder, Player) si compila anche
|
||||
su desktop con Qt 5.15 per i test:
|
||||
|
||||
```bash
|
||||
# Qt desktop (qui: Qt 5.15.2 nel kit locali)
|
||||
export PATH=<Qt5.15>/gcc_64/bin:$PATH
|
||||
mkdir -p /tmp/hh-test && cd /tmp/hh-test
|
||||
qmake /percorso/harbour-hermes/tests/core_test.pro
|
||||
make -j4
|
||||
```
|
||||
|
||||
- **Unit** (parser SSE): `./core-test`
|
||||
- **Integrazione live**: `./core-test --live`
|
||||
richiede un'istanza di prova di hermes-webui su `http://127.0.0.1:8899`
|
||||
con password in `/tmp/hermes-probe/pw` e l'audio di prova in
|
||||
`/tmp/hermes-probe/out/prova.ogg` (vedi note in `tests/core_test.cpp`).
|
||||
Il flusso coperto: login → nuova sessione → turno con streaming → TTS →
|
||||
trascrizione.
|
||||
|
||||
Nota desktop: se manca `libGL.so` di sviluppo, creare un symlink
|
||||
`libGL.so -> /usr/lib/x86_64-linux-gnu/libGL.so.1` e assicurarsi che
|
||||
`core_test.pro` punti lì (`GLSTUB`), oppure lanciare con `LD_LIBRARY_PATH`
|
||||
puntato alla lib dir di Qt.
|
||||
|
||||
## 4. Tarball per la build da Qt Creator
|
||||
|
||||
```bash
|
||||
cd /home/kaneda/workspace
|
||||
tar czf harbour-hermes-0.1.0.tar.gz \
|
||||
--transform 's,^harbour-hermes,harbour-hermes-0.1.0,' \
|
||||
--exclude='.git' harbour-hermes
|
||||
```
|
||||
|
||||
Lo spec si aspetta la directory `harbour-hermes-0.1.0/` (pattern degli altri
|
||||
progetti: cercato in modo robusto anche per i sorgenti live di sfdk).
|
||||
|
||||
## 5. Dipendenze lato server (WebUI)
|
||||
|
||||
- `/api/transcribe` richiede un provider STT attivo sul server
|
||||
(qui: faster-whisper locale — verificato con `/api/transcribe/capability`).
|
||||
- `/api/tts` con engine `edge` non richiede chiavi API.
|
||||
- Le voci italiane di edge sono state aggiunte all'allowlist del server
|
||||
(patch a `hermes-webui/api/routes.py` + test
|
||||
`tests/test_italian_voices_tts_allowlist.py`): per attivarle **riavviare
|
||||
il servizio WebUI** (`sudo systemctl restart hermes-webui`).
|
||||
- Per una dettatura italiana corretta impostare in `~/.hermes/config.yaml`:
|
||||
|
||||
```yaml
|
||||
stt:
|
||||
provider: local
|
||||
local:
|
||||
model: base
|
||||
language: it
|
||||
```
|
||||
|
||||
(senza, whisper può tradurre/maltrascrivere l'italiano). Richiede il
|
||||
riavvio del WebUI.
|
||||
@@ -0,0 +1,175 @@
|
||||
# Protocollo `hermes-webui` ↔ client
|
||||
|
||||
Documentazione del protocollo HTTP/SSE usato da harbour-hermes.
|
||||
**Validato dal vivo il 12/09/2026** contro un'istanza reale di
|
||||
`nesquena/hermes-webui` (v0.52.x, `server.py`, ThreadingHTTPServer).
|
||||
|
||||
> Il protocollo non è documentato ufficialmente: è replicato dal frontend web
|
||||
> del progetto (`static/*.js` è la reference implementation) e pinnato alla
|
||||
> versione attuale. Un aggiornamento del WebUI può richiedere aggiornamenti
|
||||
> al client.
|
||||
|
||||
Server di riferimento: in ascolto su `127.0.0.1:8787` dietro reverse proxy
|
||||
(`https://hermes.hackatoniclife.com`). TLS obbligatorio in produzione.
|
||||
|
||||
## 1. Autenticazione
|
||||
|
||||
Cookie di sessione **`hermes_session`** = `token.firma_hmac` (TTL 30 giorni
|
||||
di default; `HERMES_WEBUI_SESSION_TTL` per cambiarlo). Auth attiva solo se il
|
||||
server ha una password impostata.
|
||||
|
||||
| Endpoint | Metodo | Body / Risposta |
|
||||
|---|---|---|
|
||||
| `/api/auth/login` | POST | `{"password": "..."}` → `200 {"ok": true}` + `Set-Cookie` |
|
||||
| `/api/auth/status` | GET | `{"auth_enabled": bool, "logged_in": bool, "password_auth_enabled": bool, ...}` |
|
||||
| `/api/auth/logout` | POST | invalida il cookie |
|
||||
|
||||
- Errori: `401 {"error": "Authentication required"}`; rate limit sul login.
|
||||
- Se `auth_enabled` è `false`, tutte le API rispondono senza cookie.
|
||||
- Il client persiste SOLO il cookie (file `0600`); la password non è mai salvata.
|
||||
|
||||
## 2. Sessioni
|
||||
|
||||
`GET /api/sessions` → lista per la sidebar:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessions": [
|
||||
{
|
||||
"session_id": "171ff18f9313",
|
||||
"title": "Webui Session",
|
||||
"workspace": "/home/kaneda/.hermes/profiles/carlo/home/workspace",
|
||||
"model": "deepseek-v4-flash",
|
||||
"model_provider": "deepseek",
|
||||
"message_count": 2,
|
||||
"created_at": 1785701975.8797414,
|
||||
"updated_at": 1785701985.0872252,
|
||||
"pinned": false, "archived": false,
|
||||
"profile": "default",
|
||||
"source_tag": "webui", "session_source": "webui",
|
||||
"is_cli_session": false, "is_streaming": false,
|
||||
"active_stream_id": null
|
||||
}
|
||||
],
|
||||
"active_profile": "default",
|
||||
"server_time": 1789188804.1041577
|
||||
}
|
||||
```
|
||||
|
||||
`GET /api/session?session_id=ID` → `{"session": { ... }}` con in più:
|
||||
|
||||
- `messages`: array di `{"role": "user"|"assistant", "content": "...",
|
||||
"timestamp": 1789188869.6, ...}` (i messaggi assistente possono includere
|
||||
`reasoning`, `finish_reason`, `id`, `_turnDuration`, `_turnTps`,
|
||||
`_firstTokenMs`; i messaggi "solo tool" hanno `content` vuoto)
|
||||
- `tool_calls`, `input_tokens`, `output_tokens`, `estimated_cost`,
|
||||
`read_only`, `active_stream_id`, `enabled_toolsets`, ...
|
||||
|
||||
`POST /api/session/new` con body `{"workspace": "/path"}` (opzionale; il
|
||||
workspace deve essere sotto la home dell'utente, nella lista salvata o sotto
|
||||
il default) → `{"session": {...}}` (sessione vuota).
|
||||
|
||||
## 3. Chat con streaming
|
||||
|
||||
`POST /api/chat/start`:
|
||||
|
||||
```json
|
||||
{
|
||||
"session_id": "f517dae7a2da",
|
||||
"message": "Rispondi solo con la parola pong",
|
||||
"model": "deepseek-v4-flash",
|
||||
"model_provider": "deepseek",
|
||||
"workspace": "/home/kaneda/.hermes/profiles/carlo/home/workspace",
|
||||
"profile": "default"
|
||||
}
|
||||
```
|
||||
|
||||
Risposta:
|
||||
|
||||
```json
|
||||
{
|
||||
"stream_id": "a1e0ec82169b4673a13bf9bd466afd3f",
|
||||
"session_id": "f517dae7a2da",
|
||||
"pending_started_at": 1789188858.196835,
|
||||
"turn_id": "20260912T045418Z-ba85fcc8cebe",
|
||||
"title": "Rispondi solo con la parola pong. Non usare strumenti.",
|
||||
"effective_model_provider": "deepseek"
|
||||
}
|
||||
```
|
||||
|
||||
Errori noti: `404` sessione inesistente; `400` "Missing required field(s)";
|
||||
conflitto se la sessione ha già uno stream attivo ("session already has an
|
||||
active stream").
|
||||
|
||||
`GET /api/chat/stream?stream_id=ID` → **SSE** `text/event-stream`,
|
||||
heartbeat ogni 5 s (`: heartbeat`), ogni evento ha `id: <stream_id>:<seq>`:
|
||||
|
||||
| Evento | Payload | Uso nel client |
|
||||
|---|---|---|
|
||||
| `context_status` | `{session_id, prefill}` | ignorato |
|
||||
| `token` | `{"text": "..."}` | **delta di testo** della risposta |
|
||||
| `metering` | `{tps, ttft_ms, usage, ...}` | ignorato (v1) |
|
||||
| `done` | `{"session": {...con messages finali}, "usage": {...}}` | **risincronizza** il modello e chiude il turno |
|
||||
| `title_status` | `{session_id, status, title}` | ignorato |
|
||||
| `title` | `{session_id, "title": "..."}` | aggiorna il titolo |
|
||||
| `stream_end` | `{session_id}` | fine stream |
|
||||
|
||||
Altri eventi possono esistere (attività tool, approval/clarify) e vanno
|
||||
**ignorati** senza errore (il parser li scarta).
|
||||
|
||||
Ripresa: `GET /api/chat/stream/status?stream_id=ID` →
|
||||
`{"active": bool, "stream_id", "replay_available"}`. Se `GET /api/session`
|
||||
mostra `active_stream_id` non nullo, riagganciare lo stream con lo stesso URL.
|
||||
|
||||
Annullamento: `GET /api/chat/cancel?stream_id=ID` →
|
||||
`{"ok": true, "cancelled": bool, "stream_id"}`.
|
||||
|
||||
## 4. Voce
|
||||
|
||||
### Trascrizione (voce → testo)
|
||||
|
||||
`POST /api/transcribe` — `multipart/form-data`, campo **`file`**
|
||||
(OGG/Opus, WebM, M4A, WAV...; il server converte con ffmpeg se serve):
|
||||
|
||||
```json
|
||||
{"ok": true, "transcript": "Questa è una prova di trascrizione vocale..."}
|
||||
```
|
||||
|
||||
- Provider STT del server (qui: `local` = faster-whisper). Latenza osservata:
|
||||
~10 s per 6 s di audio (modello `base`, CPU).
|
||||
- ⚠️ **Lingua**: con `stt.local.language: ''` whisper può tradurre l'italiano
|
||||
in inglese; impostare `stt.local.language: it` in `~/.hermes/config.yaml`
|
||||
(richiede riavvio del WebUI).
|
||||
- `GET /api/transcribe/capability` → `{"ok": true, "available": true, "provider": "local"}`.
|
||||
|
||||
### Sintesi (testo → voce)
|
||||
|
||||
`POST /api/tts` — JSON:
|
||||
|
||||
```json
|
||||
{"text": "Ciao", "voice": "it-IT-ElsaNeural", "engine": "edge",
|
||||
"rate": "+0%", "pitch": "+0Hz"}
|
||||
```
|
||||
|
||||
→ `200` `audio/mpeg` (mp3) oppure `4xx` `{"error": "..."}`.
|
||||
|
||||
- Limiti server: **5000 caratteri**; rate limit **2 s per client**;
|
||||
**allowlist voci hardcoded** nella funzione `_handle_tts` di
|
||||
`api/routes.py`. Voci ammesse oggi: zh-CN (5), en-US (2), fr-CA (4),
|
||||
fr-FR (3), id-ID (1) + **it-IT (4: ElsaNeural, DiegoNeural, IsabellaNeural,
|
||||
GiuseppeMultilingualNeural)** — le italiane sono state aggiunte da noi
|
||||
(patch + test): il server va **riavviato** per attivarle.
|
||||
- Engine: `edge` (default, gratis), `elevenlabs`, `openai` (richiedono chiavi
|
||||
configurate sul server). `browser` è solo client-side, non usabile qui.
|
||||
|
||||
## 5. Note operative
|
||||
|
||||
- Tutti gli errori API sono JSON `{"error": "..."}`.
|
||||
- Il server non espone CORS: il client nativo non serve preflight.
|
||||
- Endpoint utili non usati dalla v1: `/api/clarify/*` e `/api/approval/*`
|
||||
(richieste interattive dell'agente: vanno gestite per non lasciare un run
|
||||
in attesa), `/api/upload` (allegati), `/api/settings`, `/api/models`,
|
||||
`/api/sessions/events` (SSE globale di invalidazione lista).
|
||||
- L'istanza di prova usata per la validazione era isolata (porta 8899,
|
||||
`HERMES_WEBUI_STATE_DIR` dedicata, password temporanea) e non ha toccato
|
||||
né il server live né i dati utente.
|
||||
Reference in New Issue
Block a user