AvatarPy: assistente personale con volto 3D, voce, memoria e plugin

Riscrittura in Python dell'assistente Avatar con interfaccia HUD (derivata da Mark LIV, CC BY-NC 4.0, vedi NOTICE.md).
Tre motori (Claude API, server locale OpenAI-compatibile, Claude Code), voce Kokoro/macOS, Whisper MLX,
avatar 3D con sincronizzazione labiale, memoria per categorie, allegati con OCR, monitor con avvisi,
plugin per Calendario, Mail, Promemoria, Note, Musica, app, Mac, timer, meteo, contatti, Messaggi,
file, Comandi Rapidi, browser, Telegram, WhatsApp (archivio e tempo reale).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
lucianoandClaude Fable 5.1 committed 2026-09-23 16:21:39 +02:00
commit ff79832c30
76 files changed
+82468

No files matched your search

+21
View File
@@ -0,0 +1,21 @@
# Ambienti e dipendenze
.venv/
__pycache__/
*.pyc
whatsapp_bridge/node_modules/
# Dati personali: conversazioni, archivi chat, sessioni, memoria, impostazioni e chiavi
data/
config/api_keys.json
config/settings.json
memory/long_term.json
*.session
*.session-journal
# Modelli 3D (volti personali, file grandi): aggiungili a mano in avatar3d/models/
avatar3d/models/*.glb
avatar3d/models_archivio/
# Varie
*.log
.DS_Store
+1
View File
@@ -0,0 +1 @@
3.12
+417
View File
@@ -0,0 +1,417 @@
MARK LIV — JARVIS
Copyright (c) 2026 FatihMakes
Licensed under the Creative Commons Attribution-NonCommercial 4.0
International License (CC BY-NC 4.0). Commercial use is not permitted.
Canonical text: https://creativecommons.org/licenses/by-nc/4.0/legalcode
-----------------------------------------------------------------------
Attribution-NonCommercial 4.0 International
=======================================================================
Creative Commons Corporation ("Creative Commons") is not a law firm and
does not provide legal services or legal advice. Distribution of
Creative Commons public licenses does not create a lawyer-client or
other relationship. Creative Commons makes its licenses and related
information available on an "as-is" basis. Creative Commons gives no
warranties regarding its licenses, any material licensed under their
terms and conditions, or any related information. Creative Commons
disclaims all liability for damages resulting from their use to the
fullest extent possible.
Using Creative Commons Public Licenses
Creative Commons public licenses provide a standard set of terms and
conditions that creators and other rights holders may use to share
original works of authorship and other material subject to copyright
and certain other rights specified in the public license below. The
following considerations are for informational purposes only, are not
exhaustive, and do not form part of our licenses.
Considerations for licensors: Our public licenses are
intended for use by those authorized to give the public
permission to use material in ways otherwise restricted by
copyright and certain other rights. Our licenses are
irrevocable. Licensors should read and understand the terms
and conditions of the license they choose before applying it.
Licensors should also secure all rights necessary before
applying our licenses so that the public can reuse the
material as expected. Licensors should clearly mark any
material not subject to the license. This includes other CC-
licensed material, or material used under an exception or
limitation to copyright. More considerations for licensors:
wiki.creativecommons.org/Considerations_for_licensors
Considerations for the public: By using one of our public
licenses, a licensor grants the public permission to use the
licensed material under specified terms and conditions. If
the licensor's permission is not necessary for any reason--for
example, because of any applicable exception or limitation to
copyright--then that use is not regulated by the license. Our
licenses grant only permissions under copyright and certain
other rights that a licensor has authority to grant. Use of
the licensed material may still be restricted for other
reasons, including because others have copyright or other
rights in the material. A licensor may make special requests,
such as asking that all changes be marked or described.
Although not required by our licenses, you are encouraged to
respect those requests where reasonable. More considerations
for the public:
wiki.creativecommons.org/Considerations_for_licensees
=======================================================================
Creative Commons Attribution-NonCommercial 4.0 International Public
License
By exercising the Licensed Rights (defined below), You accept and agree
to be bound by the terms and conditions of this Creative Commons
Attribution-NonCommercial 4.0 International Public License ("Public
License"). To the extent this Public License may be interpreted as a
contract, You are granted the Licensed Rights in consideration of Your
acceptance of these terms and conditions, and the Licensor grants You
such rights in consideration of benefits the Licensor receives from
making the Licensed Material available under these terms and
conditions.
Section 1 -- Definitions.
a. Adapted Material means material subject to Copyright and Similar
Rights that is derived from or based upon the Licensed Material
and in which the Licensed Material is translated, altered,
arranged, transformed, or otherwise modified in a manner requiring
permission under the Copyright and Similar Rights held by the
Licensor. For purposes of this Public License, where the Licensed
Material is a musical work, performance, or sound recording,
Adapted Material is always produced where the Licensed Material is
synched in timed relation with a moving image.
b. Adapter's License means the license You apply to Your Copyright
and Similar Rights in Your contributions to Adapted Material in
accordance with the terms and conditions of this Public License.
c. Copyright and Similar Rights means copyright and/or similar rights
closely related to copyright including, without limitation,
performance, broadcast, sound recording, and Sui Generis Database
Rights, without regard to how the rights are labeled or
categorized. For purposes of this Public License, the rights
specified in Section 2(b)(1)-(2) are not Copyright and Similar
Rights.
d. Effective Technological Measures means those measures that, in the
absence of proper authority, may not be circumvented under laws
fulfilling obligations under Article 11 of the WIPO Copyright
Treaty adopted on December 20, 1996, and/or similar international
agreements.
e. Exceptions and Limitations means fair use, fair dealing, and/or
any other exception or limitation to Copyright and Similar Rights
that applies to Your use of the Licensed Material.
f. Licensed Material means the artistic or literary work, database,
or other material to which the Licensor applied this Public
License.
g. Licensed Rights means the rights granted to You subject to the
terms and conditions of this Public License, which are limited to
all Copyright and Similar Rights that apply to Your use of the
Licensed Material and that the Licensor has authority to license.
h. Licensor means the individual(s) or entity(ies) granting rights
under this Public License.
i. NonCommercial means not primarily intended for or directed towards
commercial advantage or monetary compensation. For purposes of
this Public License, the exchange of the Licensed Material for
other material subject to Copyright and Similar Rights by digital
file-sharing or similar means is NonCommercial provided there is
no payment of monetary compensation in connection with the
exchange.
j. Share means to provide material to the public by any means or
process that requires permission under the Licensed Rights, such
as reproduction, public display, public performance, distribution,
dissemination, communication, or importation, and to make material
available to the public including in ways that members of the
public may access the material from a place and at a time
individually chosen by them.
k. Sui Generis Database Rights means rights other than copyright
resulting from Directive 96/9/EC of the European Parliament and of
the Council of 11 March 1996 on the legal protection of databases,
as amended and/or succeeded, as well as other essentially
equivalent rights anywhere in the world.
l. You means the individual or entity exercising the Licensed Rights
under this Public License. Your has a corresponding meaning.
Section 2 -- Scope.
a. License grant.
1. Subject to the terms and conditions of this Public License,
the Licensor hereby grants You a worldwide, royalty-free,
non-sublicensable, non-exclusive, irrevocable license to
exercise the Licensed Rights in the Licensed Material to:
a. reproduce and Share the Licensed Material, in whole or
in part, for NonCommercial purposes only; and
b. produce, reproduce, and Share Adapted Material for
NonCommercial purposes only.
2. Exceptions and Limitations. For the avoidance of doubt, where
Exceptions and Limitations apply to Your use, this Public
License does not apply, and You do not need to comply with
its terms and conditions.
3. Term. The term of this Public License is specified in Section
6(a).
4. Media and formats; technical modifications allowed. The
Licensor authorizes You to exercise the Licensed Rights in
all media and formats whether now known or hereafter created,
and to make technical modifications necessary to do so. The
Licensor waives and/or agrees not to assert any right or
authority to forbid You from making technical modifications
necessary to exercise the Licensed Rights, including
technical modifications necessary to circumvent Effective
Technological Measures. For purposes of this Public License,
simply making modifications authorized by this Section 2(a)
(4) never produces Adapted Material.
5. Downstream recipients.
a. Offer from the Licensor -- Licensed Material. Every
recipient of the Licensed Material automatically
receives an offer from the Licensor to exercise the
Licensed Rights under the terms and conditions of this
Public License.
b. No downstream restrictions. You may not offer or impose
any additional or different terms or conditions on, or
apply any Effective Technological Measures to, the
Licensed Material if doing so restricts exercise of the
Licensed Rights by any recipient of the Licensed
Material.
6. No endorsement. Nothing in this Public License constitutes or
may be construed as permission to assert or imply that You
are, or that Your use of the Licensed Material is, connected
with, or sponsored, endorsed, or granted official status by,
the Licensor or others designated to receive attribution as
provided in Section 3(a)(1)(A)(i).
b. Other rights.
1. Moral rights, such as the right of integrity, are not
licensed under this Public License, nor are publicity,
privacy, and/or other similar personality rights; however, to
the extent possible, the Licensor waives and/or agrees not to
assert any such rights held by the Licensor to the limited
extent necessary to allow You to exercise the Licensed
Rights, but not otherwise.
2. Patent and trademark rights are not licensed under this
Public License.
3. To the extent possible, the Licensor waives any right to
collect royalties from You for the exercise of the Licensed
Rights, whether directly or through a collecting society
under any voluntary or waivable statutory or compulsory
licensing scheme. In all other cases the Licensor expressly
reserves any right to collect such royalties, including when
the Licensed Material is used other than for NonCommercial
purposes.
Section 3 -- License Conditions.
Your exercise of the Licensed Rights is expressly made subject to the
following conditions.
a. Attribution.
1. If You Share the Licensed Material (including in modified
form), You must:
a. retain the following if it is supplied by the Licensor
with the Licensed Material:
i. identification of the creator(s) of the Licensed
Material and any others designated to receive
attribution, in any reasonable manner requested by
the Licensor (including by pseudonym if
designated);
ii. a copyright notice;
iii. a notice that refers to this Public License;
iv. a notice that refers to the disclaimer of
warranties;
v. a URI or hyperlink to the Licensed Material to the
extent reasonably practicable;
b. indicate if You modified the Licensed Material and
retain an indication of any previous modifications; and
c. indicate the Licensed Material is licensed under this
Public License, and include the text of, or the URI or
hyperlink to, this Public License.
2. You may satisfy the conditions in Section 3(a)(1) in any
reasonable manner based on the medium, means, and context in
which You Share the Licensed Material. For example, it may be
reasonable to satisfy the conditions by providing a URI or
hyperlink to a resource that includes the required
information.
3. If requested by the Licensor, You must remove any of the
information required by Section 3(a)(1)(A) to the extent
reasonably practicable.
4. If You Share Adapted Material You produce, the Adapter's
License You apply must not prevent recipients of the Adapted
Material from complying with this Public License.
Section 4 -- Sui Generis Database Rights.
Where the Licensed Rights include Sui Generis Database Rights that
apply to Your use of the Licensed Material:
a. for the avoidance of doubt, Section 2(a)(1) grants You the right
to extract, reuse, reproduce, and Share all or a substantial
portion of the contents of the database for NonCommercial purposes
only;
b. if You include all or a substantial portion of the database
contents in a database in which You have Sui Generis Database
Rights, then the database in which You have Sui Generis Database
Rights (but not its individual contents) is Adapted Material; and
c. You must comply with the conditions in Section 3(a) if You Share
all or a substantial portion of the contents of the database.
For the avoidance of doubt, this Section 4 supplements and does not
replace Your obligations under this Public License where the Licensed
Rights include other Copyright and Similar Rights.
Section 5 -- Disclaimer of Warranties and Limitation of Liability.
a. UNLESS OTHERWISE SEPARATELY UNDERTAKEN BY THE LICENSOR, TO THE
EXTENT POSSIBLE, THE LICENSOR OFFERS THE LICENSED MATERIAL AS-IS
AND AS-AVAILABLE, AND MAKES NO REPRESENTATIONS OR WARRANTIES OF
ANY KIND CONCERNING THE LICENSED MATERIAL, WHETHER EXPRESS,
IMPLIED, STATUTORY, OR OTHER. THIS INCLUDES, WITHOUT LIMITATION,
WARRANTIES OF TITLE, MERCHANTABILITY, FITNESS FOR A PARTICULAR
PURPOSE, NON-INFRINGEMENT, ABSENCE OF LATENT OR OTHER DEFECTS,
ACCURACY, OR THE PRESENCE OR ABSENCE OF ERRORS, WHETHER OR NOT
KNOWN OR DISCOVERABLE. WHERE DISCLAIMERS OF WARRANTIES ARE NOT
ALLOWED IN FULL OR IN PART, THIS DISCLAIMER MAY NOT APPLY TO YOU.
b. TO THE EXTENT POSSIBLE, IN NO EVENT WILL THE LICENSOR BE LIABLE
TO YOU ON ANY LEGAL THEORY (INCLUDING, WITHOUT LIMITATION,
NEGLIGENCE) OR OTHERWISE FOR ANY DIRECT, SPECIAL, INDIRECT,
INCIDENTAL, CONSEQUENTIAL, PUNITIVE, EXEMPLARY, OR OTHER LOSSES,
COSTS, EXPENSES, OR DAMAGES ARISING OUT OF THIS PUBLIC LICENSE OR
USE OF THE LICENSED MATERIAL, EVEN IF THE LICENSOR HAS BEEN
ADVISED OF THE POSSIBILITY OF SUCH LOSSES, COSTS, EXPENSES, OR
DAMAGES. WHERE A LIMITATION OF LIABILITY IS NOT ALLOWED IN FULL OR
IN PART, THIS LIMITATION MAY NOT APPLY TO YOU.
c. The disclaimer of warranties and limitation of liability provided
above shall be interpreted in a manner that, to the extent
possible, most closely approximates an absolute disclaimer and
waiver of all liability.
Section 6 -- Term and Termination.
a. This Public License applies for the term of the Copyright and
Similar Rights licensed here. However, if You fail to comply with
this Public License, then Your rights under this Public License
terminate automatically.
b. Where Your right to use the Licensed Material has terminated under
Section 6(a), it reinstates:
1. automatically as of the date the violation is cured, provided
it is cured within 30 days of Your discovery of the
violation; or
2. upon express reinstatement by the Licensor.
For the avoidance of doubt, this Section 6(b) does not affect any
right the Licensor may have to seek remedies for Your violations
of this Public License.
c. For the avoidance of doubt, the Licensor may also offer the
Licensed Material under separate terms or conditions or stop
distributing the Licensed Material at any time; however, doing so
will not terminate this Public License.
d. Sections 1, 5, 6, 7, and 8 survive termination of this Public
License.
Section 7 -- Other Terms and Conditions.
a. The Licensor shall not be bound by any additional or different
terms or conditions communicated by You unless expressly agreed.
b. Any arrangements, understandings, or agreements regarding the
Licensed Material not stated herein are separate from and
independent of the terms and conditions of this Public License.
Section 8 -- Interpretation.
a. For the avoidance of doubt, this Public License does not, and
shall not be interpreted to, reduce, limit, restrict, or impose
conditions on any use of the Licensed Material that could lawfully
be made without permission under this Public License.
b. To the extent possible, if any provision of this Public License is
deemed unenforceable, it shall be automatically reformed to the
minimum extent necessary to make it enforceable. If the provision
cannot be reformed, it shall be severed from this Public License
without affecting the enforceability of the remaining terms and
conditions.
c. No term or condition of this Public License will be waived and no
failure to comply consented to unless expressly agreed to by the
Licensor.
d. Nothing in this Public License constitutes or may be interpreted
as a limitation upon, or waiver of, any privileges and immunities
that apply to the Licensor or You, including from the legal
processes of any jurisdiction or authority.
=======================================================================
Creative Commons is not a party to its public
licenses. Notwithstanding, Creative Commons may elect to apply one of
its public licenses to material it publishes and in those instances
will be considered the “Licensor.” The text of the Creative Commons
public licenses is dedicated to the public domain under the CC0 Public
Domain Dedication. Except for the limited purpose of indicating that
material is shared under a Creative Commons public license or as
otherwise permitted by the Creative Commons policies published at
creativecommons.org/policies, Creative Commons does not authorize the
use of the trademark "Creative Commons" or any other trademark or logo
of Creative Commons without its prior written consent including,
without limitation, in connection with any unauthorized modifications
to any of its public licenses or any other arrangements,
understandings, or agreements concerning use of licensed material. For
the avoidance of doubt, this paragraph does not form part of the
public licenses.
Creative Commons may be contacted at creativecommons.org.
+23
View File
@@ -0,0 +1,23 @@
# Attribuzioni
L'interfaccia grafica di AvatarPy (file `ui.py`, cartella `core/` con avatar, mesh del volto,
visemi, dispositivi audio, hotkey, parola di attivazione, conferme, e la cartella `memory/`)
proviene da **Mark LIV — JARVIS**, Copyright (c) 2026 FatihMakes,
<https://github.com/FatihMakes/Mark-LIV>, rilasciato con licenza
Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0).
Il testo completo è in `LICENSE-MARK-LIV.txt`. Le funzioni di analisi audio in
`avatar/lipsync.py` sono adattate da `main.py` dello stesso progetto.
Modifiche rispetto all'originale: rimossa la dipendenza da Gemini, aggiunti i pulsanti
"Motore e Voce" e "Nuova conversazione", cambiato l'elenco delle voci nel pannello Customise,
caratteri ingranditi, aggiunta la vista avatar 3D. Su richiesta dell'utente i riferimenti a
"MARK LIV" e "FatihMakes" sono stati tolti dall'interfaccia (23 settembre 2026): l'attribuzione
richiesta dalla licenza resta in questo file e nel README.
**Questa licenza vieta l'uso commerciale** dell'app nel suo insieme finché contiene questi file.
Il modello del volto (`core/face_model.obj`) deriva dal canonical face model di MediaPipe
(Apache 2.0). Kokoro-82M è di hexgrad (Apache 2.0). Whisper è di OpenAI (MIT), nella
versione MLX di mlx-community.
Il resto del codice (cartella `avatar/`, `main.py`) è opera di questo progetto.
+145
View File
@@ -0,0 +1,145 @@
# AvatarPy — assistente personale con volto, in Python
Riscrittura in Python dell'assistente "Avatar": stessa idea (un assistente che ascolta, parla, ricorda e cerca sul web) con l'interfaccia HUD di **Mark LIV**: un volto tridimensionale renderizzato in software, con sincronizzazione labiale reale, forma d'onda, pannello di log e impostazioni.
L'interfaccia è riusata con licenza CC BY-NC 4.0 (uso personale, non commerciale): vedi `NOTICE.md`.
## Cosa fa
- **Conversazione a mani libere**: il microfono è sempre in ascolto; quando smetti di parlare, la frase viene trascrizione in locale (Whisper via MLX) e inviata al motore. In alternativa: premi-per-parlare, parola di attivazione "Hey Jarvis" (openwakeword, installabile dal pannello), o il campo di testo.
- **Voce**: Kokoro in locale con le voci italiane Sara e Nicola (PyTorch), oppure la voce di sistema di macOS. Le frasi vengono sintetizzate in anticipo mentre l'app pronuncia la precedente.
- **Volto**: la bocca segue lo spettro dell'audio (50 forme al secondo) fuso con il testo pronunciato; lo sguardo e le sopracciglia seguono lo stato (ascolta, pensa, parla, dorme).
- **Tre motori**, selezionabili dal pulsante "Motore & Voce":
- *Claude (Anthropic)*: Claude Opus 5 via API, con ricerca web server-side. Serve una chiave, salvata nel portachiavi di macOS.
- *Server locale compatibile OpenAI*: vLLM, Ollama, LM Studio. Ricerca web con una chiave Brave Search (opzionale).
- *Claude Code*: usa Claude Code installato sul Mac e il suo accesso, senza chiave. Tre livelli: solo conversazione, lettura dei file, completo.
- **Memoria a lungo termine** per categorie (identità, preferenze, progetti, relazioni, desideri, note) nel file `memory/long_term.json`, condivisa da tutti i motori e visibile dal pulsante "Memory" del pannello. Con Claude Code la memoria passa da un piccolo server MCP incluso.
## Avatar 3D personale
Il pulsante "HUD" nel pannello alterna tre viste: volto animato, nucleo, **avatar 3D**. La terza mostra un modello GLB (`avatar3d/models/avatar.glb`) renderizzato con three.js dentro l'interfaccia: inquadratura sul busto (clic per la figura intera), animazione di riposo, testa e occhi che guardano la camera e seguono lo stato (ascolta, pensa, parla, dorme), sbattito delle palpebre, forma d'onda e stato in basso.
La **sincronizzazione labiale** usa i blend shape del modello: visemi Oculus (`viseme_aa`, `viseme_O`, `viseme_I`…) e ARKit (`jawOpen`, `mouthSmileLeft`, `mouthFunnel`, `eyeBlinkLeft`…), pilotati dalle forme della bocca calcolate dall'audio. I modelli inclusi sono avatar Avaturn esportati con i blend shape e le ossa degli occhi: `allegra_2.glb`, `allegra.glb` e `luza.glb`. Si sceglie dal menu "Avatar 3D" in "Motore e Voce", che elenca i file GLB presenti in `avatar3d/models`; il cambio è immediato. Per aggiungere un avatar basta copiare il suo GLB in quella cartella: il nome del file diventa il nome nel menu.
Per usare un altro modello sostituisci il GLB: vanno bene avatar Avaturn, Ready Player Me o Mixamo con scheletro standard (ossa `Head`, `Neck`, `Hips`, opzionali `LeftEye`/`RightEye`). Senza blend shape la bocca resta ferma e si muove solo la testa.
## Allegati
Trascina un file nella zona "File upload" (o clicca per sceglierlo) e poi fai la domanda: il file viene passato al motore con il messaggio successivo, una volta sola.
- **Testo, PDF, Word, RTF**: viene allegato il testo estratto (fino a 8000 caratteri).
- **Immagini** (PNG, JPEG, HEIC…): il testo presente viene riconosciuto in locale con l'OCR di macOS, quindi funziona con tutti i motori, anche quello locale. Con Claude API e Claude Code l'immagine viene anche vista davvero, per descrizioni e domande sul contenuto.
## Importare le memorie da ChatGPT (o da un altro assistente)
1. In ChatGPT chiedi: "Elenca tutte le memorie che hai salvato su di me, una per riga" e copia la risposta in un file di testo (oppure copia l'elenco da Impostazioni › Personalizzazione › Gestisci memorie).
2. Nell'app trascina il file nella zona "File upload" e di': "importa queste memorie". L'assistente le salva nella categoria giusta con lo strumento `salva_memorie`, in una volta sola.
3. Controlla il risultato dal pulsante "Memory" e cancella ciò che non vuoi tenere.
## Monitor con avvisi (Mail, WhatsApp, Telegram)
Attivabile in "Motore e Voce" › Monitor. Ogni minuto (intervallo regolabile) l'app raccoglie i messaggi nuovi da Mail (posta in arrivo), WhatsApp (ponte in tempo reale) e Telegram (account collegato), li passa a un modello veloce con le regole che scrivi tu, e per quelli che meritano attenzione crea un avviso: riga rossa nel log, notifica di macOS e, se vuoi, annuncio a voce. Le regole predefinite segnalano richieste di aiuto o assistenza, domande che aspettano risposta, problemi, urgenze, scadenze, pagamenti e appuntamenti da confermare, e ignorano newsletter, promozioni e notifiche automatiche.
A voce: "ci sono avvisi?", "chiudi l'avviso di Marco", "controlla adesso". Il primo giro dopo l'attivazione prende solo la linea di base: vengono valutati i messaggi arrivati da quel momento in poi. Con il motore Claude Code la valutazione usa Haiku a basso sforzo; con Claude API usa Haiku 4.5; con il server locale usa il modello locale.
## Plugin: comandare il Mac a voce
I plugin sono file Python nella cartella `plugins/`: ognuno dichiara uno o più strumenti (nome, descrizione, parametri, funzione) e l'assistente li usa da tutti e tre i motori. Il pannello "Plugins" li elenca e permette di disattivarli.
Incluso: **`calendario`**, che legge, crea, sposta e cancella eventi del Calendario di macOS tramite EventKit. Esempi a voce: "che impegni ho domani", "sono libero venerdì pomeriggio", "segna dentista giovedì alle 15", "sposta la riunione di lunedì alle 11", "cancella il pranzo di mercoledì". La cancellazione chiede conferma sullo schermo (pulsante Confirm nel pannello); con Claude Code l'assistente chiede conferma a voce. Al primo uso macOS chiede il permesso per il Calendario.
Incluso anche **`mail`**, per l'app Mail di macOS: "ho email nuove?", "cerca le email di Marco", "leggimi l'ultima email di Amazon", "scrivi a mario@esempio.it che arrivo alle dieci". Elenco, ricerca e lettura usano l'indice locale di Mail (istantanei, anche a Mail chiusa); l'invio passa da Mail con conferma sullo schermo e richiede il permesso Automazione al primo uso.
Altri plugin inclusi, tutti locali:
| Plugin | Cosa fa | Esempi |
|---|---|---|
| `promemoria` | app Promemoria (EventKit) | "ricordami di chiamare Marco alle 17", "aggiungi latte alla spesa", "cosa devo fare oggi" |
| `note` | app Note | "prendi nota: …", "leggimi la nota sul progetto", "aggiungi alla nota Idee…" |
| `musica` | Apple Music (controlli anche su Spotify) | "metti la playlist Mattina", "pausa", "prossimo", "cosa sta suonando", "volume musica al 30" |
| `app` | apri/chiudi app, siti e file | "apri Safari sul Corriere", "avvia Xcode", "chiudi Mail", "che app sono aperte" |
| `mac` | volume, luminosità, Non disturbare, Wi-Fi, stato, blocco schermo, screenshot | "alza il volume", "spegni il Wi-Fi", "quanta batteria ho", "blocca il Mac" |
| `timer` | timer, sveglie, attività programmate anche giornaliere | "timer di 10 minuti per la pasta", "svegliami alle 7", "ogni mattina alle 8 leggimi la posta" |
| `meteo` | previsioni (Open-Meteo, senza chiave) | "che tempo fa domani a Milano" |
| `contatti` | Contatti | "qual è il numero di Anna", "l'email di Luca" |
| `messaggi` | iMessage con conferma | "manda un messaggio a Luca che arrivo alle 9" |
| `file` | Spotlight, cartelle, lettura txt/PDF/Word | "cerca il PDF del contratto", "cosa c'è sul Desktop", "riassumi il documento X" |
| `comandi_rapidi` | Comandi Rapidi, anche HomeKit | "accendi le luci del soggiorno" (se esiste il comando rapido) |
| `browser` | legge pagine, cerca su Google/YouTube/Maps/Amazon/Wikipedia | "leggimi questa pagina", "metti su YouTube i Pink Floyd" |
| `buongiorno` | riepilogo: impegni, promemoria, email, meteo | "buongiorno", "come si presenta la giornata" |
| `whatsapp_archivio` | chat WhatsApp esportate: importazione, ricerca full-text, lettura per periodo, statistiche | "cosa mi ha detto Marco sulla riunione?", "riassumi la chat con Anna dell'ultima settimana", "quando abbiamo parlato del dentista?" |
| `whatsapp_live` | WhatsApp in tempo reale come dispositivo collegato (Baileys): novità, invio con conferma, annunci vocali | "ci sono novità su WhatsApp?", "chi mi ha scritto?", "scrivi a Marco su WhatsApp che arrivo" |
| `telegram` | Telegram con il tuo account (Telethon): non letti, lettura, ricerca, invio con conferma | "ho messaggi su Telegram?", "leggimi la chat con Anna", "scrivi a Luca su Telegram che sono in ritardo" |
**WhatsApp in tempo reale** usa un client non ufficiale (Baileys, in `whatsapp_bridge/`, Node.js) collegato come "dispositivo collegato". È contro i termini di WhatsApp e Meta può bloccare il numero: l'app non invia mai senza conferma e non fa azioni di massa, ma il rischio resta a carico di chi lo attiva. Per collegarlo: "Motore e Voce" › WhatsApp › "Collega (QR)", poi sul telefono WhatsApp › Impostazioni › Dispositivi collegati › Collega un dispositivo. Da quel momento i messaggi in arrivo (e quelli che invii da altri dispositivi) finiscono nell'archivio, in tempo reale; con "Annuncia a voce" l'assistente li legge appena arrivano. Il ponte parte da solo all'avvio dell'app una volta collegato. Richiede Node.js (`npm install` in `whatsapp_bridge/` è già fatto).
Per lo storico, l'app lavora sulle chat **esportate**: in WhatsApp apri la chat › Altro › Esporta chat (senza media), poi di' all'assistente "importa la chat WhatsApp che è in Download" (o "importa tutte le chat della cartella X") oppure copia il file `.txt` o lo `.zip` in `data/whatsapp/`. Gli zip "WhatsApp Chat - Nome.zip" prendono il nome del contatto. L'indice è locale e istantaneo anche con decine di migliaia di messaggi; riesportando la chat, l'archivio si aggiorna.
Telegram richiede una configurazione una tantum in "Motore e Voce": api id e api hash creati su my.telegram.org (API development tools), poi il numero di telefono, "Invia codice", il codice ricevuto su Telegram e "Accedi" (più la password se hai la verifica in due passaggi). La sessione resta salvata in `data/telegram.session`; "Esci" la cancella. WhatsApp non ha un'API per account personali e non è supportato.
Le attività programmate (plugin `timer`) vengono eseguite dall'app quando è aperta: un avviso viene detto a voce, un comando viene eseguito come se lo avessi chiesto tu. Non disturbare richiede due Comandi Rapidi chiamati "Non disturbare ON/OFF". Ogni app controllata chiede il permesso Automazione la prima volta.
Per scriverne uno nuovo copia la struttura di `plugins/calendario.py`: una lista `TOOLS` con `name`, `description`, `parameters` (JSON Schema) e `run(parameters, ctx)`; `ctx` offre `say` (parlare), `log` e `confirm` (conferma a schermo per azioni irreversibili). Riavvia l'app dopo aver aggiunto il file. Funziona anche il formato a singolo `PLUGIN` di Mark-LIV.
## Requisiti
- macOS su Apple Silicon, Python 3.12, [uv](https://docs.astral.sh/uv/).
- `brew install espeak-ng portaudio` (fonetica italiana per Kokoro e audio).
- Permessi macOS richiesti al primo uso: Microfono, Calendario.
- Per il motore Claude Code: Claude Code installato e collegato.
## Avvio
```bash
cd AvatarPy
uv sync
uv run python main.py
```
Al primo avvio si apre la finestra "Motore & Voce" se manca la configurazione. La prima volta vengono scaricati i modelli di Kokoro (circa 330 MB) e di Whisper (circa 500 MB per "small"); poi tutto resta in locale.
## Uso
| Azione | Come |
|---|---|
| Parlare | Parla e basta: la frase parte quando fai una pausa di circa un secondo |
| Premi-per-parlare | Pulsante "Push-to-talk" nel cassetto rapido; il microfono si apre solo con il tasto premuto |
| Scrivere | Campo di testo in basso a destra, Invio per inviare |
| Interrompere | Esc, oppure il pulsante di interruzione |
| Silenziare il microfono | F4 |
| Vedere o cancellare la memoria | Pulsante "Memory" |
| Cambiare nome, colore e voce | Pulsante "Customise assistant" |
| Cambiare motore, chiavi, riconoscimento vocale | Pulsante "Motore & Voce" |
| Ricominciare da zero | Pulsante "Nuova conversazione" |
## Profili di Claude Code
Se sul Mac ci sono più profili (`~/.claude`, `~/.claude-<nome>`), scegli quello giusto in "Motore & Voce"; "Verifica" mostra l'account collegato. Un errore `401 OAuth access token is invalid` significa che l'accesso di quel profilo è scaduto: `CLAUDE_CONFIG_DIR=<cartella> claude auth login`.
## Struttura
```
main.py avvio, collegamento tra interfaccia e assistente
ui.py, core/, memory/ interfaccia e moduli da Mark LIV (vedi NOTICE.md)
avatar/
assistant.py orchestratore: microfono → Whisper → motore → voce, stati e log
audio.py microfono con rilevamento della voce; riproduzione con visemi
lipsync.py livello audio e forme della bocca dallo spettro
stt.py Whisper (MLX, ripiego su faster-whisper)
tts.py Kokoro e voce di sistema, spezzettamento in frasi
memory_tools.py strumenti di memoria comuni ai motori
memory_mcp.py server MCP che espone la memoria a Claude Code
websearch.py ricerca Brave (motore locale)
settings.py impostazioni, chiavi nel portachiavi
settings_dialog.py finestra "Motore e Voce"
avatar3d.py vista 3D (QWebEngineView + server locale) sincronizzata con stati e visemi
persona.md personalità (modificabile)
plugins.py registro dei plugin (strumenti per i motori e per MCP)
engines/ anthropic_engine.py, openai_compat.py, claude_code.py
plugins/ plugin (calendario.py incluso)
avatar3d/ pagina three.js, modello GLB e librerie
config/ api_keys.json (identità), settings.json (impostazioni)
data/ cronologie delle conversazioni per motore
```
La personalità dell'assistente è in `avatar/persona.md`; il nome e il tuo nome si impostano dal pulsante "Customise assistant".
View File
Whitespace-only changes.
+503
View File
@@ -0,0 +1,503 @@
"""Orchestratore: microfono → trascrizione → motore → voce, con stati e log sulla UI."""
from __future__ import annotations
import json
import queue
import threading
import time
import numpy as np
from memory import config_manager as cm
from .audio import Microphone, Player
from .engines.anthropic_engine import AnthropicEngine
from .engines.claude_code import ClaudeCodeEngine
from .engines.openai_compat import OpenAICompatEngine
from .settings import CONFIG_DIR, Settings
from .stt import Transcriber
from .tts import SentenceSplitter, clean_for_speech, make_voice
END = object()
STATUS_LABELS = {
"thinking": "THINKING", "searching": "PROCESSING", "memory": "PROCESSING",
"working": "PROCESSING", "responding": "THINKING",
}
# Frasi che Whisper "inventa" su silenzio o rumore di fondo (sottotitoli visti in addestramento).
_HALLUCINATIONS = ("sottotitoli", "a cura di", "grazie per aver guardato", "iscriviti al canale",
"www.", "amara.org", "subtitles", "thank you for watching")
def _is_hallucination(text: str) -> bool:
t = text.lower()
return any(h in t for h in _HALLUCINATIONS)
def _identity() -> tuple[str, str]:
try:
d = json.loads((CONFIG_DIR / "api_keys.json").read_text(encoding="utf-8"))
return (d.get("assistant_name") or "Ava").strip(), (d.get("user_name") or "").strip()
except Exception:
return "Ava", ""
class Assistant:
def __init__(self, ui, settings: Settings) -> None:
self.ui = ui
self.settings = settings
self.name, self.user_name = _identity()
self._abort = threading.Event()
self._turn_lock = threading.Lock()
self._busy = False
self._speaking = False
self._tail_until = 0.0
self._turn = 0
self._ptt_enabled = False
self._ptt_held = False
self._wake_enabled = False
self._awake = True
self._wake = None
self._last_speech = time.monotonic()
self._engine = None
self._engine_key = None
self._last_attachment: tuple[str, float] | None = None
self._speech_q: queue.Queue = queue.Queue()
self._audio_q: queue.Queue = queue.Queue(maxsize=3)
self.player = Player(self.ui.push_visemes, self.ui.set_audio_level)
self.mic = Microphone(self._on_utterance, self._mic_level, self._can_listen,
float(settings.get("vad_threshold", 0.08)), on_frame=self._on_mic_frame)
self.stt = Transcriber(str(settings.get("stt_model")), on_status=self._note)
self.voice = make_voice(settings, on_status=self._note)
# ── Avvio ────────────────────────────────────────────────────────────
def start(self) -> None:
threading.Thread(target=self._synth_loop, name="tts-synth", daemon=True).start()
threading.Thread(target=self._play_loop, name="tts-play", daemon=True).start()
threading.Thread(target=self._warmup, name="warmup", daemon=True).start()
threading.Thread(target=self._scheduler_loop, name="scheduler", daemon=True).start()
def _start_whatsapp_bridge(self) -> None:
try:
from avatar import whatsapp_bridge as wb
if self.settings.get("whatsapp_live") or wb.is_linked():
self.ui.write_log(f"SYS: WhatsApp — {wb.start(self.user_name)}")
self._wa_last_check = time.strftime("%Y-%m-%d %H:%M:%S")
except Exception as err:
self.ui.write_log(f"ERR: WhatsApp — {err}")
def _announce_whatsapp(self) -> None:
"""Legge i messaggi WhatsApp arrivati dall'ultimo controllo e, se richiesto, li annuncia."""
import sqlite3
from avatar.settings import DATA_DIR
db = DATA_DIR / "whatsapp" / "index.sqlite"
if not db.exists():
return
since = getattr(self, "_wa_last_check", None) or time.strftime("%Y-%m-%d %H:%M:%S")
now = time.strftime("%Y-%m-%d %H:%M:%S")
con = sqlite3.connect(f"file:{db}?mode=ro", uri=True)
try:
rows = con.execute("SELECT ts, chat, sender, text FROM messages WHERE source = 'live' AND from_me = 0 AND ts > ? ORDER BY ts LIMIT 5", (since,)).fetchall()
finally:
con.close()
self._wa_last_check = now
for ts, chat, sender, text in rows:
who = chat if sender == chat else f"{sender} nel gruppo {chat}"
self.ui.write_log(f"WhatsApp: {who}: {text[:200]}")
if self.settings.get("whatsapp_annuncia") and not self._busy and not self._speaking:
self.say(f"Messaggio WhatsApp da {who}: {text[:220]}")
def _scheduler_loop(self) -> None:
"""Fa scattare timer, sveglie e attività programmate dal plugin timer."""
import importlib
while True:
time.sleep(5)
try:
mod = importlib.import_module("avatar_plugins.timer")
except Exception:
try:
from avatar.plugins import registry
registry.load()
mod = importlib.import_module("avatar_plugins.timer")
except Exception:
continue
try:
self._announce_whatsapp()
except Exception as err:
print(f"[whatsapp] {err}")
try:
for item in mod.due():
testo = str(item.get("testo", ""))
if item.get("comando"):
self.ui.write_log(f"SYS: Attività programmata: {testo}")
self.handle_text(testo)
else:
self.ui.write_log(f"SYS: Timer: {testo}")
self.say(f"Promemoria: {testo}")
except Exception as err:
print(f"[scheduler] {err}")
def _warmup(self) -> None:
self.ui.set_state("PROCESSING")
self.ui.write_log(f"SYS: {self.name} si sta avviando: carico voce e riconoscimento vocale…")
try:
self.voice.load()
except Exception as err:
self.ui.write_log(f"ERR: Voce non disponibile — {err}")
try:
self.stt.load()
except Exception as err:
self.ui.write_log(f"ERR: Riconoscimento vocale non disponibile — {err}")
self.reopen_audio()
self._ensure_engine()
self._start_whatsapp_bridge()
try:
from avatar import monitor as _mon
_mon.monitor = _mon.Monitor(self.ui, self.say, self.settings)
_mon.monitor.start()
except Exception as err:
self.ui.write_log(f"ERR: Monitor — {err}")
self.ui.set_state("LISTENING")
self.ui.write_log(f"SYS: {self.name} è pronta. Parla o scrivi.")
def reopen_audio(self) -> None:
try:
self.player.open(cm.get_output_device())
except Exception as err:
self.ui.write_log(f"ERR: Uscita audio — {err}")
try:
self.mic.start(cm.get_input_device())
except Exception as err:
self.ui.write_log(f"ERR: Microfono — {err}")
def _note(self, msg: str | None) -> None:
if msg:
self.ui.write_log(f"SYS: {msg}")
# ── Motore ───────────────────────────────────────────────────────────
def _ensure_engine(self):
s = self.settings
provider = s.get("provider")
self.name, self.user_name = _identity()
key = (provider, s.get("effort"), s.get("local_base_url"), s.get("local_model"), s.get("search_api_key"),
s.get("claudecode_model"), s.get("claudecode_access"), s.get("claudecode_config_dir"),
s.get("claudecode_path"), self.name, self.user_name, s.get_secret("anthropic_api_key")[-6:],
s.get_secret("local_api_key")[-4:])
if self._engine and self._engine_key == key:
return self._engine
self._engine, self._engine_key = None, key
if provider == "local":
if s.get("local_base_url") and s.get("local_model"):
self._engine = OpenAICompatEngine(s.get("local_base_url"), s.get("local_model"), s.get_secret("local_api_key"),
s.get("search_api_key"), self.name, self.user_name)
elif provider == "claudecode":
self._engine = ClaudeCodeEngine(s.get("claudecode_model") or "sonnet", s.get("claudecode_access") or "chat",
s.get("claudecode_path") or "", s.get("claudecode_config_dir") or "",
self.name, self.user_name, s.get("effort"))
else:
api_key = s.get_secret("anthropic_api_key")
if api_key:
self._engine = AnthropicEngine(api_key, self.name, self.user_name, s.get("effort"))
return self._engine
def reconfigure(self) -> None:
"""Dopo un salvataggio delle impostazioni."""
self.interrupt(silent=True)
self._ensure_engine()
self.reload_voice()
model = str(self.settings.get("stt_model"))
if model != self.stt.model:
self.stt = Transcriber(model, on_status=self._note)
threading.Thread(target=self.stt.load, daemon=True).start()
self.mic.threshold = float(self.settings.get("vad_threshold", 0.08))
try:
view = getattr(self.ui._win, "avatar3d", None)
if view is not None:
view.set_model(str(self.settings.get("avatar_model") or ""))
except Exception as err:
self.ui.write_log(f"ERR: Avatar 3D — {err}")
self.ui.write_log(f"SYS: Impostazioni applicate — motore {self.settings.get('provider')}.")
def reload_voice(self, from_customise: bool = False) -> None:
"""Ricarica la voce. `from_customise`: la scelta arriva dal pannello Customise
(Sara / Nicola / Sistema) e va copiata nelle impostazioni; altrimenti comandano
le impostazioni di "Motore e Voce" e il pannello viene allineato."""
mapping = {"Sara": ("kokoro", "if_sara"), "Nicola": ("kokoro", "im_nicola"), "Sistema": ("system", None)}
if from_customise:
chosen = (cm.get_voice() or "").strip()
if chosen in mapping:
engine, voice = mapping[chosen]
self.settings.set("tts_engine", engine)
if voice:
self.settings.set("kokoro_voice", voice)
self.settings.save()
else:
label = "Sistema" if self.settings.get("tts_engine") == "system" else \
{"if_sara": "Sara", "im_nicola": "Nicola"}.get(self.settings.get("kokoro_voice"), "Sara")
try:
if cm.get_voice() != label:
cm.save_voice(label)
except Exception:
pass
new = make_voice(self.settings, on_status=self._note)
if type(new) is type(self.voice) and getattr(new, "voice", None) == getattr(self.voice, "voice", None):
return
self.voice = new
self.ui.write_log(f"SYS: Voce: {getattr(new, 'voice', '') or 'sistema'} ({'Kokoro' if self.settings.get('tts_engine') == 'kokoro' else 'macOS'}).")
threading.Thread(target=self._safe_load_voice, daemon=True).start()
def _safe_load_voice(self) -> None:
try:
self.voice.load()
except Exception as err:
self.ui.write_log(f"ERR: Voce non disponibile — {err}")
def reset_conversation(self) -> None:
self.interrupt(silent=True)
if self._engine:
self._engine.reset()
self.ui.write_log("SYS: Nuova conversazione.")
# ── Ascolto ──────────────────────────────────────────────────────────
def _can_listen(self) -> bool:
if self.ui.muted or self._busy or self._speaking or not self._awake:
return False
if time.monotonic() < self._tail_until:
return False
if self._ptt_enabled and not self._ptt_held:
return False
return True
def _mic_level(self, level: float) -> None:
self.ui.set_audio_level(level)
def _on_mic_frame(self, frame: np.ndarray) -> None:
if self._wake is not None and self._wake_enabled and not self._awake:
try:
self._wake.feed(frame)
except Exception:
pass
if self._wake_enabled and self._awake and not self._busy and not self._speaking:
if time.monotonic() - self._last_speech > 120:
self._sleep("silenzio")
def _on_utterance(self, audio: np.ndarray) -> None:
self._last_speech = time.monotonic()
self.ui.set_state("PROCESSING")
try:
text = self.stt.transcribe(audio)
except Exception as err:
self.ui.write_log(f"ERR: Trascrizione — {err}")
self.ui.set_state("LISTENING")
return
text = text.strip()
if len(text) < 2 or _is_hallucination(text):
self.ui.set_state("LISTENING")
return
self.ui.write_log(f"You: {text}")
self.handle_text(text)
# ── Turno di conversazione ───────────────────────────────────────────
def handle_text(self, text: str) -> None:
"""Chiamabile da qualunque thread (UI, microfono, quiz…)."""
text = (text or "").strip()
if not text:
return
if self._busy or self._speaking:
self.interrupt(silent=True)
threading.Thread(target=self._run_turn, args=(text,), daemon=True).start()
def say(self, text: str) -> None:
"""Pronuncia un testo senza passare dal motore."""
for s in SentenceSplitter().push(text + "\n") + []:
self._speech_q.put((self._turn, s))
self._speech_q.put((self._turn, END))
def _run_turn(self, text: str) -> None:
with self._turn_lock:
engine = self._ensure_engine()
if engine is None:
self.ui.write_log("ERR: Motore non configurato: apri Motore & Voce e completa le impostazioni.")
self.ui.set_state("LISTENING")
return
self._abort.clear()
self._busy = True
self._turn += 1
turn = self._turn
self.ui.set_state("THINKING")
splitter = SentenceSplitter()
def emit(ev: dict) -> None:
if turn != self._turn:
return
t = ev.get("type")
if t == "status":
if not self._speaking:
self.ui.set_state(STATUS_LABELS.get(ev["status"], "THINKING"))
if ev["status"] == "searching":
self.ui.write_log("SYS: Cerco sul web…")
elif ev["status"] == "working":
self.ui.write_log(f"SYS: Lavoro sul Mac ({ev.get('detail', '')})…")
elif t == "text":
for s in splitter.push(ev["delta"]):
self._speech_q.put((turn, s))
elif t == "memory_saved":
self.ui.write_log(f"SYS: Memoria salvata — {ev['text']}")
elif t == "memory_removed":
self.ui.write_log("SYS: Memoria cancellata.")
elif t == "done":
for s in splitter.flush():
self._speech_q.put((turn, s))
self.ui.write_log(f"{self.name}: {ev['text']}")
if ev.get("sources"):
self.ui.show_content("Fonti", "\n".join(f"• {s['title']}\n {s['url']}" for s in ev["sources"]))
elif t == "error":
if not ev.get("aborted"):
self.ui.write_log(f"ERR: {ev['message']}")
self._speech_q.put((turn, clean_for_speech("Scusa, c'è stato un problema: " + ev["message"])[:300]))
# Allegato dalla zona "File upload" (una volta per file)
engine_text, image, attach_path = text, None, ""
try:
path = self.ui.current_file
if path:
from pathlib import Path as _P
p = _P(path)
key = (str(p), p.stat().st_mtime)
if p.exists() and key != self._last_attachment:
self._last_attachment = key
from avatar.attachments import describe
self.ui.write_log(f"FILE: {p.name}")
ctx_text, image = describe(p)
engine_text = f"{text}\n\n{ctx_text}"
attach_path = str(p)
except Exception as err:
self.ui.write_log(f"ERR: Allegato — {err}")
try:
engine.send(engine_text, emit, self._abort, image=image, attach_path=attach_path)
finally:
self._busy = False
self._speech_q.put((turn, END))
# ── Voce: sintesi in anticipo e riproduzione ─────────────────────────
def _synth_loop(self) -> None:
while True:
turn, item = self._speech_q.get()
if turn != self._turn:
continue
if item is END:
self._audio_q.put((turn, END, None))
continue
try:
audio = self.voice.synthesize(item)
except Exception as err:
self.ui.write_log(f"ERR: Sintesi vocale — {err}")
continue
if turn == self._turn:
self._audio_q.put((turn, item, audio))
def _play_loop(self) -> None:
while True:
turn, text, audio = self._audio_q.get()
if turn != self._turn:
continue
if text is END:
self.player.drain()
self._speaking = False
self._tail_until = time.monotonic() + 0.5
if not self._busy and turn == self._turn:
self.ui.set_state("LISTENING" if self._awake else "SLEEPING")
continue
if audio is None or len(audio) == 0:
continue
self._speaking = True
self.ui.set_state("SPEAKING")
self.player.play(audio, text, self._abort)
self._last_speech = time.monotonic()
def interrupt(self, silent: bool = False) -> None:
self._abort.set()
self._turn += 1
for q in (self._speech_q, self._audio_q):
try:
while True:
q.get_nowait()
except queue.Empty:
pass
self.player.abort()
self._speaking = False
self._busy = False
self._tail_until = time.monotonic() + 0.4
if not silent:
self.ui.write_log("SYS: Interrotta. Ti ascolto.")
self.ui.set_state("LISTENING" if self._awake else "SLEEPING")
# ── Push-to-talk e wake word ─────────────────────────────────────────
def set_push_to_talk(self, enabled: bool) -> str:
self._ptt_enabled = bool(enabled)
self._ptt_held = False
return "window"
def ptt_hold(self, held: bool) -> None:
self._ptt_held = bool(held)
if held and not self._awake:
self._wake_up("tasto")
self.ui.set_state("LISTENING" if held else ("LISTENING" if not self._ptt_enabled else "SLEEPING"))
def wake_get_state(self) -> dict:
ready = False
try:
from core.wake_word import is_ready
ready = is_ready()
except Exception:
pass
return {"enabled": self._wake_enabled, "awake": self._awake, "ready": ready}
def on_wake_toggle(self, enable: bool) -> str:
if enable:
try:
from core.wake_word import WakeWordDetector, is_ready
if not is_ready():
return "Il modello della parola di attivazione non è installato."
self._wake = WakeWordDetector(on_detect=lambda: self._wake_up("parola di attivazione"))
if not self._wake.start():
self._wake = None
return "Impossibile avviare la parola di attivazione."
except Exception as err:
self._wake = None
return f"Parola di attivazione non disponibile: {err}"
self._wake_enabled = True
self._last_speech = time.monotonic()
self.ui.write_log("SYS: Parola di attivazione attiva: di' «Hey Jarvis» per svegliarmi.")
return "on"
self._wake_enabled = False
if self._wake:
try:
self._wake.stop()
except Exception:
pass
self._wake = None
self._wake_up("disattivata")
return "off"
def on_wake_manual(self) -> None:
if self._awake:
self._sleep("manuale")
else:
self._wake_up("manuale")
def _sleep(self, reason: str) -> None:
if not self._awake:
return
self._awake = False
self.ui.set_state("SLEEPING")
self.ui.write_log(f"SYS: In pausa ({reason}).")
def _wake_up(self, reason: str) -> None:
self._last_speech = time.monotonic()
if self._awake:
return
self._awake = True
self.ui.set_state("LISTENING")
self.ui.write_log(f"SYS: Sveglia ({reason}).")
+74
View File
@@ -0,0 +1,74 @@
"""Allegati dalla zona "File upload": testo dai documenti, OCR e immagine dalle foto."""
from __future__ import annotations
import io
import re
from pathlib import Path
IMAGE_EXT = {".png", ".jpg", ".jpeg", ".gif", ".webp", ".heic", ".tiff", ".bmp"}
MAX_TEXT = 8000
def ocr(path: Path) -> str:
"""Testo riconosciuto nell'immagine con il framework Vision di macOS."""
import Quartz
import Vision
from Foundation import NSURL
src = Quartz.CGImageSourceCreateWithURL(NSURL.fileURLWithPath_(str(path)), None)
if src is None:
return ""
img = Quartz.CGImageSourceCreateImageAtIndex(src, 0, None)
if img is None:
return ""
req = Vision.VNRecognizeTextRequest.alloc().init()
req.setRecognitionLevel_(Vision.VNRequestTextRecognitionLevelAccurate)
req.setRecognitionLanguages_(["it-IT", "en-US"])
req.setUsesLanguageCorrection_(True)
handler = Vision.VNImageRequestHandler.alloc().initWithCGImage_options_(img, None)
ok, err = handler.performRequests_error_([req], None)
if not ok:
return ""
lines = []
for obs in req.results() or []:
cands = obs.topCandidates_(1)
if cands:
lines.append(str(cands[0].string()))
return "\n".join(lines)
def image_payload(path: Path, max_side: int = 1568) -> tuple[str, bytes]:
"""Immagine ridotta per i modelli con visione: (media_type, bytes)."""
from PIL import Image
im = Image.open(path)
im = im.convert("RGB")
im.thumbnail((max_side, max_side))
buf = io.BytesIO()
im.save(buf, format="JPEG", quality=85)
return "image/jpeg", buf.getvalue()
def describe(path: Path) -> tuple[str, tuple[str, bytes] | None]:
"""Restituisce (contesto testuale da aggiungere al messaggio, immagine per la visione o None)."""
ext = path.suffix.lower()
if ext in IMAGE_EXT:
text = ""
try:
text = ocr(path)
except Exception as err:
text = f"(OCR non riuscito: {err})"
image = None
try:
image = image_payload(path)
except Exception:
pass
ctx = f"[Allegato immagine: {path.name}]\nTesto riconosciuto nell'immagine (OCR):\n{text.strip() or '(nessun testo)'}"
return ctx, image
# Documenti: riusa il lettore del plugin file
try:
from avatar.plugins import registry
registry.load()
text = registry.run("file_leggi", {"percorso": str(path)})
except Exception as err:
text = f"(lettura non riuscita: {err})"
text = re.sub(r"\s+", " ", text)[:MAX_TEXT]
return f"[Allegato: {path.name}]\n{text}", None
+241
View File
@@ -0,0 +1,241 @@
"""Microfono con rilevamento della voce e riproduzione con sincronizzazione labiale."""
from __future__ import annotations
import queue
import threading
import time
from collections import deque
from typing import Callable
import numpy as np
import sounddevice as sd
from core import audio_devices
from core.viseme import VisemeStream
from .lipsync import HOP_SECONDS, float_to_pcm_scale, pcm_level, pcm_visemes
MIC_RATE = 16000
OUT_RATE = 24000
MIC_BLOCK = 480 # 30 ms
PRE_ROLL_BLOCKS = 10 # 300 ms prima dell'inizio del parlato
START_BLOCKS = 3 # 90 ms di voce per iniziare
END_SILENCE_S = 0.8
MIN_SPEECH_S = 0.35
MAX_UTTERANCE_S = 30.0
class Microphone:
"""Cattura continua; segmenta le frasi con una soglia di volume adattiva."""
def __init__(
self,
on_utterance: Callable[[np.ndarray], None],
on_level: Callable[[float], None],
can_listen: Callable[[], bool],
threshold: float = 0.08,
on_frame: Callable[[np.ndarray], None] | None = None,
) -> None:
self._on_utterance = on_utterance
self._on_level = on_level
self._can_listen = can_listen
self._on_frame = on_frame
self.threshold = threshold
self._q: queue.Queue[np.ndarray] = queue.Queue()
self._stream: sd.InputStream | None = None
self._thread: threading.Thread | None = None
self._running = False
self._device_name = ""
def start(self, device_name: str = "") -> None:
self.stop()
self._device_name = device_name
try:
device = audio_devices.resolve(device_name, "input") if device_name else None
except Exception:
device = None
self._stream = sd.InputStream(
samplerate=MIC_RATE, channels=1, dtype="int16", blocksize=MIC_BLOCK,
device=device, callback=self._callback,
)
self._stream.start()
self._running = True
self._thread = threading.Thread(target=self._loop, name="mic-vad", daemon=True)
self._thread.start()
def stop(self) -> None:
self._running = False
if self._stream is not None:
try:
self._stream.stop()
self._stream.close()
except Exception:
pass
self._stream = None
def _callback(self, indata, frames, t, status) -> None:
self._q.put(indata[:, 0].copy())
def _loop(self) -> None:
pre = deque(maxlen=PRE_ROLL_BLOCKS)
collecting: list[np.ndarray] = []
voiced_run = 0
silence_s = 0.0
speech_s = 0.0
block_s = MIC_BLOCK / MIC_RATE
floor = 0.0
while self._running:
try:
frame = self._q.get(timeout=0.5)
except queue.Empty:
continue
if self._on_frame:
try:
self._on_frame(frame)
except Exception:
pass
level = pcm_level(frame)
# Rumore di fondo: media lenta dei livelli bassi.
if level < self.threshold:
floor = floor * 0.98 + level * 0.02
if not self._can_listen():
# Se stavamo raccogliendo (per esempio il tasto premi-per-parlare
# è stato rilasciato) chiudi la frase invece di perderla.
if collecting and speech_s >= MIN_SPEECH_S:
audio = np.concatenate(collecting)
collecting = []
try:
self._on_utterance(audio)
except Exception:
pass
pre.clear()
collecting = []
voiced_run = 0
continue
self._on_level(level)
voiced = level > max(self.threshold, floor * 2.5 + 0.02)
if not collecting:
pre.append(frame)
voiced_run = voiced_run + 1 if voiced else 0
if voiced_run >= START_BLOCKS:
collecting = list(pre)
silence_s = 0.0
speech_s = voiced_run * block_s
voiced_run = 0
continue
collecting.append(frame)
if voiced:
silence_s = 0.0
speech_s += block_s
else:
silence_s += block_s
total_s = len(collecting) * block_s
if silence_s >= END_SILENCE_S or total_s >= MAX_UTTERANCE_S:
audio = np.concatenate(collecting)
collecting = []
pre.clear()
if speech_s >= MIN_SPEECH_S:
try:
self._on_utterance(audio)
except Exception:
pass
class Player:
"""Riproduce audio a 24 kHz e programma le forme della bocca sul tempo reale."""
def __init__(self, on_visemes: Callable, on_level: Callable[[float], None]) -> None:
self._on_visemes = on_visemes
self._on_level = on_level
self._visemes = VisemeStream()
self._stream: sd.OutputStream | None = None
self._lock = threading.Lock()
self._device_name = ""
self._cursor = 0.0
def open(self, device_name: str = "") -> None:
self.close()
self._device_name = device_name
try:
device = audio_devices.resolve(device_name, "output") if device_name else None
except Exception:
device = None
self._stream = sd.OutputStream(samplerate=OUT_RATE, channels=1, dtype="float32", device=device)
self._stream.start()
def close(self) -> None:
if self._stream is not None:
try:
self._stream.stop()
self._stream.close()
except Exception:
pass
self._stream = None
@property
def latency(self) -> float:
try:
return float(self._stream.latency) if self._stream else 0.05
except Exception:
return 0.05
def play(self, audio: np.ndarray, text: str, stop: threading.Event) -> None:
"""Bloccante: suona `audio` (float32 mono 24 kHz) mentre anima la bocca."""
if self._stream is None:
self.open(self._device_name)
assert self._stream is not None
audio = np.asarray(audio, dtype=np.float32).reshape(-1)
if audio.size == 0:
return
with self._lock:
try:
self._visemes.feed_text(text)
except Exception:
pass
# Il cursore è il momento in cui il prossimo blocco inizierà a suonare.
now = time.time() + self.latency + 0.04
if self._cursor < now:
self._cursor = now
block = int(OUT_RATE * 0.2)
for i in range(0, audio.size, block):
if stop.is_set():
break
chunk = audio[i:i + block]
frames = pcm_visemes(float_to_pcm_scale(chunk), OUT_RATE)
if frames:
try:
frames = self._visemes.frames(frames, HOP_SECONDS)
except Exception:
pass
self._on_visemes(frames, HOP_SECONDS, self._cursor)
self._on_level(max(f[0] for f in frames))
else:
self._on_level(pcm_level(float_to_pcm_scale(chunk)))
self._cursor += chunk.size / OUT_RATE
try:
self._stream.write(chunk.reshape(-1, 1))
except Exception:
break
if stop.is_set():
self.abort()
def abort(self) -> None:
"""Ferma subito, scartando l'audio in coda."""
try:
self._visemes.reset()
except Exception:
pass
self._cursor = 0.0
if self._stream is not None:
try:
self._stream.abort()
self._stream.start()
except Exception:
self.open(self._device_name)
self._on_level(0.0)
def drain(self) -> None:
"""Aspetta che l'audio scritto abbia finito di suonare."""
wait = max(0.0, self._cursor - time.time())
if wait > 0:
time.sleep(min(wait, 5.0))
self._on_level(0.0)
+131
View File
@@ -0,0 +1,131 @@
"""Vista 3D dell'avatar (three.js in un QWebEngineView), sincronizzata con stati, volume e visemi."""
from __future__ import annotations
import http.server
import threading
import time
from functools import partial
from pathlib import Path
from PyQt6.QtCore import QTimer, QUrl
from PyQt6.QtGui import QColor
from PyQt6.QtWebEngineCore import QWebEnginePage
from PyQt6.QtWebEngineWidgets import QWebEngineView
ROOT = Path(__file__).resolve().parent.parent / "avatar3d"
MODELS_DIR = ROOT / "models"
def list_models() -> list[str]:
"""Nomi (senza estensione) dei modelli GLB disponibili."""
return sorted(p.stem for p in MODELS_DIR.glob("*.glb"))
class _Quiet(http.server.SimpleHTTPRequestHandler):
def log_message(self, *args) -> None:
pass
_server_port: int | None = None
def serve_root() -> int:
"""Server HTTP locale (una volta sola): i moduli ES non si caricano da file://."""
global _server_port
if _server_port:
return _server_port
handler = partial(_Quiet, directory=str(ROOT))
srv = http.server.ThreadingHTTPServer(("127.0.0.1", 0), handler)
threading.Thread(target=srv.serve_forever, name="avatar3d-http", daemon=True).start()
_server_port = srv.server_address[1]
return _server_port
class _Page(QWebEnginePage):
def javaScriptConsoleMessage(self, level, message, line, source) -> None: # noqa: N802
print(f"[avatar3d] {message} ({Path(source).name}:{line})")
class Avatar3DView(QWebEngineView):
def __init__(self, parent=None) -> None:
super().__init__(parent)
self._page = _Page(self)
self.setPage(self._page)
self._page.setBackgroundColor(QColor("#00060a"))
self._state = "LISTENING"
self._level = 0.0
self._sched: tuple[list, float, float] | None = None
self._ready = False
self._model = ""
self.loadFinished.connect(self._on_loaded)
try:
from avatar.settings import Settings
model = str(Settings().get("avatar_model") or "")
except Exception:
model = ""
self.set_model(model)
self._timer = QTimer(self)
self._timer.timeout.connect(self._tick)
self._timer.start(33)
def _on_loaded(self, ok: bool) -> None:
self._ready = ok
def set_model(self, name: str) -> None:
"""Carica (o ricarica) la pagina con il modello indicato."""
available = list_models()
if name not in available:
name = available[0] if available else name
if name == self._model and self._ready:
return
self._model = name
self._ready = False
self.load(QUrl(f"http://127.0.0.1:{serve_root()}/index.html?model={name}"))
# ── API thread-safe (valori semplici, letti dal timer sul thread Qt) ──
def set_state(self, state: str) -> None:
self._state = str(state or "LISTENING")
def set_audio_level(self, level: float) -> None:
try:
self._level = max(self._level, min(1.0, max(0.0, float(level))))
except (TypeError, ValueError):
pass
def push_visemes(self, frames, hop: float, at: float) -> None:
if not frames:
return
new = list(frames)
cur = self._sched
if cur is not None and abs(cur[2] - hop) < 1e-6:
old, t0, _ = cur
i = int(round((at - t0) / hop))
if 0 <= i <= len(old) + 1:
merged = old[:i] + new
played = int((time.time() - t0) / hop) - 2
if played > 60:
merged, t0 = merged[played:], t0 + played * hop
self._sched = (merged, t0, hop)
return
self._sched = (new, float(at), float(hop))
def glance(self, dx: float, dy: float, hold: float = 1.1) -> None:
if self._ready:
self._page.runJavaScript(f"window.avatar3d && avatar3d.glance({dx:.3f},{dy:.3f},{hold:.2f})")
def _tick(self) -> None:
open_, width, vis_level = 0.0, 0.0, 0.0
sched = self._sched
if sched is not None:
frames, t0, hop = sched
i = int((time.time() - t0) / hop)
if i >= len(frames):
self._sched = None
elif i >= 0:
vis_level, open_, width = frames[i]
level = max(self._level, vis_level)
self._level *= 0.82
if self._ready:
self._page.runJavaScript(
f"window.avatar3d && avatar3d.update({level:.3f},{open_:.3f},{width:.3f},'{self._state}')"
)
View File
Whitespace-only changes.
+169
View File
@@ -0,0 +1,169 @@
"""Motore Claude (API Anthropic) con ricerca web server-side e strumenti di memoria."""
from __future__ import annotations
import threading
import anthropic
from avatar.memory_tools import anthropic_tools, memory_prompt, parse_args, run_tool
from avatar.plugins import registry
from .base import Emit, History, persona_text, today_label, user_block
MODEL = "claude-opus-5"
MAX_ROUNDS = 12
WEB_SEARCH = {
"type": "web_search_20260209",
"name": "web_search",
"max_uses": 5,
"user_location": {"type": "approximate", "country": "IT", "timezone": "Europe/Rome"},
}
class Aborted(Exception):
pass
def friendly_error(err: Exception) -> str:
if isinstance(err, anthropic.AuthenticationError):
return "La chiave API Anthropic non è valida. Controllala nelle impostazioni."
if isinstance(err, anthropic.PermissionDeniedError):
return "La chiave API non ha i permessi per questo modello."
if isinstance(err, anthropic.RateLimitError):
return "Troppe richieste in poco tempo. Riprova tra qualche secondo."
if isinstance(err, anthropic.BadRequestError):
return f"Richiesta non valida: {err.message}"
if isinstance(err, anthropic.APIConnectionError):
return "Non riesco a raggiungere il servizio Anthropic. Controlla la connessione."
if isinstance(err, anthropic.APIStatusError):
return f"Errore del servizio ({err.status_code}): {err.message}"
return str(err)
class AnthropicEngine:
name = "anthropic"
def __init__(self, api_key: str, assistant_name: str, user_name: str, effort: str) -> None:
self.client = anthropic.Anthropic(api_key=api_key)
self.assistant_name = assistant_name
self.user_name = user_name
self.effort = effort
self.history = History("anthropic")
self._compaction = True
def reset(self) -> None:
self.history.clear()
def _system(self) -> list[dict]:
return [
{"type": "text", "text": persona_text(self.assistant_name), "cache_control": {"type": "ephemeral"}},
{"type": "text", "text": user_block(self.user_name, memory_prompt())},
{"type": "text", "text": f"Adesso è {today_label()}.\n\nSe uno strumento risponde con [CONFIRMATION_PENDING], chiedi all'utente di confermare sul pannello e non dire che è fatto."},
]
@staticmethod
def _sources(content: list[dict]) -> list[dict]:
seen: dict[str, dict] = {}
for block in content:
for c in block.get("citations") or []:
if c.get("type") == "web_search_result_location" and c.get("url") and c["url"] not in seen:
seen[c["url"]] = {"title": c.get("title") or c["url"], "url": c["url"]}
return list(seen.values())
def send(self, user_text: str, emit: Emit, abort: threading.Event, image: tuple[str, bytes] | None = None, **_) -> None:
if image:
import base64
media, data = image
self.history.messages.append({"role": "user", "content": [
{"type": "image", "source": {"type": "base64", "media_type": media, "data": base64.b64encode(data).decode()}},
{"type": "text", "text": user_text},
]})
else:
self.history.messages.append({"role": "user", "content": user_text})
full = ""
sources: list[dict] = []
try:
emit({"type": "status", "status": "thinking"})
for _ in range(MAX_ROUNDS):
kwargs: dict = dict(
model=MODEL,
max_tokens=16000,
betas=["server-side-fallback-2026-07-01"] + (["compact-2026-01-12"] if self._compaction else []),
fallbacks="default",
output_config={"effort": self.effort},
system=self._system(),
tools=[WEB_SEARCH, *anthropic_tools(), *registry.anthropic_tools()],
messages=self.history.messages,
)
if self._compaction:
kwargs["context_management"] = {"edits": [{"type": "compact_20260112"}]}
announced = False
try:
with self.client.beta.messages.stream(**kwargs) as stream:
for event in stream:
if abort.is_set():
raise Aborted()
if event.type == "content_block_start":
t = event.content_block.type
if t == "server_tool_use":
emit({"type": "status", "status": "searching"})
elif t == "tool_use":
emit({"type": "status", "status": "memory"})
elif t == "thinking":
emit({"type": "status", "status": "thinking"})
elif event.type == "content_block_delta" and event.delta.type == "text_delta":
if not announced:
announced = True
emit({"type": "status", "status": "responding"})
full += event.delta.text
emit({"type": "text", "delta": event.delta.text})
message = stream.get_final_message()
except anthropic.BadRequestError as err:
if self._compaction and not full:
print(f"[Claude] compattazione rifiutata, la disattivo: {err.message}")
self._compaction = False
continue
raise
content = message.model_dump(mode="json")["content"]
sources += self._sources(content)
self.history.messages.append({"role": "assistant", "content": content})
if message.stop_reason == "refusal":
if not full:
why = getattr(message.stop_details, "explanation", None) if message.stop_details else None
full = "Non posso aiutarti con questa richiesta." + (f" ({why})" if why else "")
break
if message.stop_reason == "pause_turn":
continue
tool_uses = [b for b in message.content if b.type == "tool_use"]
if message.stop_reason != "tool_use" or not tool_uses:
break
results = []
for b in tool_uses:
if registry.has(b.name):
emit({"type": "status", "status": "working", "detail": b.name})
res, ev = registry.run(b.name, parse_args(b.input)), None
else:
res, ev = run_tool(b.name, parse_args(b.input))
if ev:
emit(ev)
results.append({"type": "tool_result", "tool_use_id": b.id, "content": res})
self.history.messages.append({"role": "user", "content": results})
emit({"type": "status", "status": "thinking"})
full = full.strip() or "(nessuna risposta)"
self.history.save()
emit({"type": "done", "text": full, "sources": sources})
except Aborted:
self._rollback()
emit({"type": "error", "message": "Risposta interrotta.", "aborted": True})
except Exception as err:
self._rollback()
emit({"type": "error", "message": friendly_error(err)})
def _rollback(self) -> None:
msgs = self.history.messages
if msgs and msgs[-1].get("role") == "user" and (isinstance(msgs[-1].get("content"), str) or
(isinstance(msgs[-1].get("content"), list) and any(b.get("type") == "image" for b in msgs[-1]["content"]))):
msgs.pop()
self.history.save()
+69
View File
@@ -0,0 +1,69 @@
"""Tipi comuni ai motori di conversazione."""
from __future__ import annotations
import json
import locale
import threading
from datetime import datetime
from pathlib import Path
from typing import Any, Callable, Protocol
from avatar.settings import BASE_DIR, DATA_DIR
Event = dict[str, Any]
Emit = Callable[[Event], None]
PERSONA_FILE = BASE_DIR / "avatar" / "persona.md"
class ChatBackend(Protocol):
def send(self, user_text: str, emit: Emit, abort: threading.Event) -> None: ...
def reset(self) -> None: ...
def today_label() -> str:
try:
locale.setlocale(locale.LC_TIME, "it_IT.UTF-8")
except Exception:
pass
d = datetime.now()
giorni = ["lunedì", "martedì", "mercoledì", "giovedì", "venerdì", "sabato", "domenica"]
mesi = ["gennaio", "febbraio", "marzo", "aprile", "maggio", "giugno", "luglio",
"agosto", "settembre", "ottobre", "novembre", "dicembre"]
return f"{giorni[d.weekday()]} {d.day} {mesi[d.month - 1]} {d.year}, ore {d:%H:%M}"
def persona_text(assistant_name: str) -> str:
try:
text = PERSONA_FILE.read_text(encoding="utf-8")
except Exception:
text = "Sei Ava, un'assistente personale gentile e concreta. Rispondi in italiano."
return text.replace("Ava", assistant_name or "Ava")
def user_block(user_name: str, memory_block: str) -> str:
who = f"L'utente si chiama {user_name}." if user_name else "Non conosci ancora il nome dell'utente: chiediglielo con naturalezza e salvalo."
return f"{who}\n\nMemoria a lungo termine sull'utente:\n{memory_block}"
class History:
"""Cronologia su disco per un motore."""
def __init__(self, name: str) -> None:
self.file: Path = DATA_DIR / f"conversation-{name}.json"
self.messages: list[Any] = []
self.meta: dict[str, Any] = {}
try:
d = json.loads(self.file.read_text(encoding="utf-8"))
self.messages = list(d.get("messages", []))
self.meta = dict(d.get("meta", {}))
except Exception:
pass
def save(self) -> None:
self.file.parent.mkdir(parents=True, exist_ok=True)
self.file.write_text(json.dumps({"messages": self.messages, "meta": self.meta}, ensure_ascii=False), encoding="utf-8")
def clear(self) -> None:
self.messages, self.meta = [], {}
self.save()
+217
View File
@@ -0,0 +1,217 @@
"""Motore Claude Code: avvia `claude -p` in streaming JSON con un server MCP per la memoria."""
from __future__ import annotations
import json
import os
import shutil
import subprocess
import sys
import threading
import uuid
from pathlib import Path
from memory import memory_manager as mm
from avatar.memory_tools import memory_prompt
from .base import Emit, History, persona_text, today_label, user_block
MEMORY_TOOLS = ["mcp__avatar"] # tutti gli strumenti del server MCP dell'app (memoria e plugin)
ACCESS = {
"chat": {"tools": "WebSearch,WebFetch", "allowed": ["WebSearch", "WebFetch"], "mode": None},
"read": {"tools": "Read,Glob,Grep,WebSearch,WebFetch", "allowed": ["Read", "Glob", "Grep", "WebSearch", "WebFetch"], "mode": None},
"full": {"tools": "default",
"allowed": ["Bash", "Read", "Edit", "Write", "MultiEdit", "NotebookEdit", "Glob", "Grep", "WebSearch", "WebFetch", "Agent"],
"mode": "acceptEdits"},
}
CANDIDATES = [Path.home() / ".local/bin/claude", Path("/opt/homebrew/bin/claude"), Path("/usr/local/bin/claude"),
Path.home() / ".claude/local/claude", Path.home() / ".npm-global/bin/claude"]
MCP_SCRIPT = Path(__file__).resolve().parent.parent / "memory_mcp.py"
def resolve_claude(custom: str = "") -> str | None:
if custom.strip():
return custom.strip() if Path(custom.strip()).exists() else None
found = shutil.which("claude")
if found:
return found
for p in CANDIDATES:
if p.exists():
return str(p)
return None
def claude_env(config_dir: str) -> dict:
env = dict(os.environ)
env.pop("CLAUDECODE", None)
if config_dir.strip():
env["CLAUDE_CONFIG_DIR"] = config_dir.strip()
return env
def check_claude_code(custom_path: str, config_dir: str) -> dict:
"""Comando, profilo e account collegato."""
binary = resolve_claude(custom_path)
home = Path.home()
profiles = sorted(str(p) for p in home.iterdir() if p.is_dir() and (p.name == ".claude" or p.name.startswith(".claude-")))
effective = config_dir.strip() or os.environ.get("CLAUDE_CONFIG_DIR") or str(home / ".claude")
st = {"binary": binary, "configDir": effective, "loggedIn": False, "profiles": profiles}
if not binary:
st["error"] = "Comando claude non trovato."
return st
try:
out = subprocess.run([binary, "auth", "status"], env=claude_env(config_dir), capture_output=True, text=True, timeout=15)
d = json.loads(out.stdout)
st.update(loggedIn=bool(d.get("loggedIn")), email=d.get("email"), authMethod=d.get("authMethod"),
configDir=d.get("configDirectory") or effective)
except Exception as err:
st["error"] = str(err)[:200]
return st
class ClaudeCodeEngine:
name = "claudecode"
def __init__(self, model: str, access: str, claude_path: str, config_dir: str,
assistant_name: str, user_name: str, effort: str) -> None:
self.model, self.access, self.claude_path, self.config_dir = model, access, claude_path, config_dir
self.assistant_name, self.user_name, self.effort = assistant_name, user_name, effort
self.history = History("claudecode")
self._proc: subprocess.Popen | None = None
def reset(self) -> None:
self.history.clear()
def _system_prompt(self) -> str:
access = {
"chat": "Non hai accesso ai file o ai comandi del Mac: se ti chiedono di farlo, spiega che serve alzare il livello di accesso nelle impostazioni dell'app.",
"read": "Puoi leggere file e cercare nelle cartelle dell'utente, ma non modificare nulla né eseguire comandi.",
"full": "Puoi leggere e modificare file ed eseguire comandi sul Mac dell'utente. Fallo con prudenza e spiega brevemente cosa hai fatto.",
}[self.access]
return "\n\n".join([
persona_text(self.assistant_name),
user_block(self.user_name, memory_prompt()),
f"Adesso è {today_label()}.",
"Strumenti di memoria (obbligatori): per ricordare un fatto sull'utente chiama `mcp__avatar__salva_memoria`; per cancellarne uno `mcp__avatar__dimentica_memoria`; per cercare `mcp__avatar__cerca_memoria`. Non dire mai di aver salvato qualcosa senza averlo chiamato davvero. Gli altri strumenti `mcp__avatar__*` (per esempio il calendario) agiscono sul Mac dell'utente: usali quando servono, e per le azioni irreversibili chiedi conferma a voce prima di richiamarli con confermato=true.",
access,
"Stai parlando dentro un'app vocale: rispondi in modo conversazionale, senza intestazioni Markdown né tabelle, e non citare nomi di file o strumenti interni a meno che non serva.",
])
def _args(self, session_id: str, resume: bool, attach_path: str = "") -> list[str]:
flags = dict(ACCESS[self.access])
if attach_path:
# Permesso di lettura limitato al file allegato (immagini incluse: Claude Code le vede).
if "Read" not in flags["tools"] and flags["tools"] != "default":
flags["tools"] = flags["tools"] + ",Read"
flags["allowed"] = [*flags["allowed"], f"Read(//{attach_path.lstrip('/')})"]
mcp = {"mcpServers": {"avatar": {"command": sys.executable, "args": [str(MCP_SCRIPT)]}}}
args = ["-p", "--output-format", "stream-json", "--verbose", "--include-partial-messages",
"--model", self.model, "--effort", self.effort,
"--tools", flags["tools"], "--allowedTools", ",".join(flags["allowed"] + MEMORY_TOOLS),
"--mcp-config", json.dumps(mcp), "--strict-mcp-config",
"--system-prompt-snapshot", "off",
"--resume" if resume else "--session-id", session_id]
if flags["mode"]:
args += ["--permission-mode", flags["mode"]]
args += ["--system-prompt" if self.access == "chat" else "--append-system-prompt", self._system_prompt()]
return args
def _run_once(self, text: str, session_id: str, resume: bool, emit: Emit, abort: threading.Event, attach_path: str = "") -> dict:
binary = resolve_claude(self.claude_path)
if not binary:
raise RuntimeError("Non trovo il comando `claude`. Installa Claude Code o indica il percorso nelle impostazioni.")
proc = subprocess.Popen([binary, *self._args(session_id, resume, attach_path)], cwd=str(Path.home()), env=claude_env(self.config_dir),
stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True)
self._proc = proc
out = {"text": "", "initialized": False, "error": None, "result": ""}
announced = False
def watch_abort():
abort.wait()
if proc.poll() is None:
proc.terminate()
threading.Thread(target=watch_abort, daemon=True).start()
try:
proc.stdin.write(text)
proc.stdin.close()
except Exception:
pass
for line in proc.stdout:
line = line.strip()
if not line:
continue
try:
msg = json.loads(line)
except Exception:
continue
t = msg.get("type")
if t == "system" and msg.get("subtype") == "init":
out["initialized"] = True
elif t == "stream_event":
ev = msg.get("event") or {}
if ev.get("type") == "content_block_start" and (ev.get("content_block") or {}).get("type") == "tool_use":
name = ev["content_block"].get("name", "")
if name.startswith("mcp__memoria__"):
emit({"type": "status", "status": "memory"})
elif name in ("WebSearch", "WebFetch"):
emit({"type": "status", "status": "searching"})
else:
emit({"type": "status", "status": "working", "detail": name})
elif ev.get("type") == "content_block_delta" and (ev.get("delta") or {}).get("type") == "text_delta":
delta = ev["delta"].get("text", "")
if delta:
if not announced:
announced = True
emit({"type": "status", "status": "responding"})
out["text"] += delta
emit({"type": "text", "delta": delta})
elif ev.get("type") == "message_start" and announced:
emit({"type": "status", "status": "thinking"})
elif t == "result":
out["result"] = msg.get("result") or ""
if msg.get("is_error") or (msg.get("subtype") and msg["subtype"] != "success"):
errs = msg.get("errors") or []
out["error"] = "; ".join(map(str, errs)) or out["result"] or msg.get("subtype")
stderr = proc.stderr.read() if proc.stderr else ""
out["code"] = proc.wait()
out["stderr"] = stderr
self._proc = None
if not out["text"] and out["result"] and not out["error"]:
out["text"] = out["result"]
emit({"type": "text", "delta": out["result"]})
return out
def send(self, user_text: str, emit: Emit, abort: threading.Event, attach_path: str = "", **_) -> None:
if attach_path:
user_text += f"\n\n(Il file allegato è in {attach_path}: se ti serve vederlo, leggilo con lo strumento Read.)"
before = {(r["category"], r["key"]) for r in mm.all_entries_for_ui()}
session_id = str(self.history.meta.get("sessionId") or "")
resume = bool(session_id)
session_id = session_id or str(uuid.uuid4())
try:
emit({"type": "status", "status": "thinking"})
out = self._run_once(user_text, session_id, resume, emit, abort, attach_path)
if not abort.is_set() and resume and not out["initialized"] and out["code"] != 0:
print("[Claude Code] sessione non ripresa, ne creo una nuova:", out["stderr"][:200])
session_id, resume = str(uuid.uuid4()), False
out = self._run_once(user_text, session_id, resume, emit, abort, attach_path)
if out["initialized"]:
self.history.meta["sessionId"] = session_id
for r in mm.all_entries_for_ui():
if (r["category"], r["key"]) not in before:
emit({"type": "memory_saved", "text": r["value"]})
if abort.is_set():
self.history.save()
emit({"type": "error", "message": "Risposta interrotta.", "aborted": True})
return
if out["error"] or (out["code"] != 0 and not out["text"]):
detail = out["error"] or " ".join(out["stderr"].strip().splitlines()[-3:])
raise RuntimeError(detail or f"Claude Code è terminato con codice {out['code']}")
text = out["text"].strip() or "(nessuna risposta)"
self.history.messages.append({"role": "user", "text": user_text})
self.history.messages.append({"role": "assistant", "text": text})
self.history.save()
emit({"type": "done", "text": text, "sources": []})
except Exception as err:
self.history.save()
emit({"type": "error", "message": str(err)})
+220
View File
@@ -0,0 +1,220 @@
"""Motore per server compatibili OpenAI (vLLM, Ollama, LM Studio…)."""
from __future__ import annotations
import re
import threading
import openai
from avatar.memory_tools import memory_prompt, openai_tools, parse_args, run_tool
from avatar.plugins import registry
from avatar.websearch import brave_search
from .base import Emit, History, persona_text, today_label, user_block
MAX_ROUNDS = 8
MAX_HISTORY = 30
SEARCH_TOOL = {
"type": "function",
"function": {
"name": "cerca_web",
"description": "Cerca sul web informazioni aggiornate (notizie, prezzi, orari, meteo, eventi). Restituisce titoli, link e descrizioni.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]},
},
}
class Aborted(Exception):
pass
def friendly_error(err: Exception, base_url: str) -> str:
if isinstance(err, openai.APIConnectionError):
return f"Non riesco a raggiungere il server locale su {base_url}. È avviato?"
if isinstance(err, openai.AuthenticationError):
return "Il server locale ha rifiutato la chiave API."
if isinstance(err, openai.NotFoundError):
return "Modello non trovato sul server locale. Controlla il nome nelle impostazioni."
if isinstance(err, openai.BadRequestError):
return f"Il server locale ha rifiutato la richiesta: {err.message}"
if isinstance(err, openai.APIStatusError):
return f"Errore del server locale ({err.status_code}): {err.message}"
return str(err)
class ThinkFilter:
"""Nasconde i blocchi <think>…</think> dei modelli ragionanti."""
def __init__(self) -> None:
self.inside = False
self.carry = ""
def push(self, delta: str) -> str:
text, self.carry, out = self.carry + delta, "", ""
while text:
if self.inside:
end = text.find("</think>")
if end < 0:
self.carry = text[-8:]
return out
text = text[end + 8:].lstrip()
self.inside = False
else:
start = text.find("<think>")
if start < 0:
m = re.search(r"<(t(h(i(n(k)?)?)?)?)?$", text)
if m:
self.carry = m.group(0)
out += text[: m.start()]
else:
out += text
return out
out += text[:start]
text = text[start + 7:]
self.inside = True
return out
class OpenAICompatEngine:
name = "local"
def __init__(self, base_url: str, model: str, api_key: str, search_api_key: str,
assistant_name: str, user_name: str) -> None:
self.base_url, self.model, self.search_api_key = base_url, model, search_api_key
self.assistant_name, self.user_name = assistant_name, user_name
self.client = openai.OpenAI(base_url=base_url, api_key=api_key or "non-necessaria")
self.history = History("local")
self._tools_supported = True
def reset(self) -> None:
self.history.clear()
def _system(self) -> dict:
if self._tools_supported:
note = ("\n\nRegole sugli strumenti (obbligatorie):\n"
"- Quando l'utente ti chiede di ricordare qualcosa, o ti dice un fatto importante su di sé, DEVI chiamare salva_memoria prima di rispondere. Non dire mai di aver salvato senza averlo chiamato davvero.\n"
"- Per cancellare una memoria chiama dimentica_memoria; per cercarne una non presente nel prompt chiama cerca_memoria.\n"
"- Per azioni sul Mac (per esempio il calendario) usa gli strumenti dedicati. Se uno risponde con [CONFIRMATION_PENDING], chiedi all'utente di confermare sul pannello e non dire che è fatto.")
note += ("\n- Per informazioni aggiornate chiama cerca_web e rispondi in base ai risultati." if self.search_api_key
else "\n- Non hai accesso al web: se ti chiedono informazioni aggiornate, dillo chiaramente.")
else:
note = "\n\nNota: in questa modalità non hai strumenti (niente memoria automatica né ricerca web)."
return {"role": "system", "content": f"{persona_text(self.assistant_name)}\n\n{user_block(self.user_name, memory_prompt())}\n\nAdesso è {today_label()}.{note}"}
def _recent(self) -> list:
msgs = self.history.messages
if len(msgs) <= MAX_HISTORY:
return msgs
start = len(msgs) - MAX_HISTORY
while start < len(msgs) and msgs[start].get("role") != "user":
start += 1
return msgs[start:]
def _tools(self):
if not self._tools_supported:
return None
return openai_tools() + registry.openai_tools() + ([SEARCH_TOOL] if self.search_api_key else [])
def _run_tool(self, name: str, raw_args: str, emit: Emit, sources: list) -> str:
args = parse_args(raw_args)
if name == "cerca_web":
emit({"type": "status", "status": "searching"})
try:
hits = brave_search(str(args.get("query", "")), self.search_api_key)
except Exception as err:
return f"Errore nella ricerca: {err}"
for h in hits[:5]:
if all(s["url"] != h["url"] for s in sources):
sources.append({"title": h["title"], "url": h["url"]})
return "\n".join(f"{i+1}. {h['title']}\n {h['url']}\n {h['description']}" for i, h in enumerate(hits)) or "Nessun risultato."
if registry.has(name):
emit({"type": "status", "status": "working", "detail": name})
return registry.run(name, args)
emit({"type": "status", "status": "memory"})
res, ev = run_tool(name, args)
if ev:
emit(ev)
return res
def send(self, user_text: str, emit: Emit, abort: threading.Event, **_) -> None:
self.history.messages.append({"role": "user", "content": user_text})
full = ""
sources: list[dict] = []
try:
emit({"type": "status", "status": "thinking"})
for rnd in range(MAX_ROUNDS):
tools = self._tools()
kwargs: dict = dict(model=self.model, messages=[self._system(), *self._recent()], stream=True)
if tools:
kwargs.update(tools=tools, tool_choice="auto")
try:
stream = self.client.chat.completions.create(**kwargs)
except openai.BadRequestError as err:
if tools and not full:
print(f"[locale] il server rifiuta gli strumenti, li disattivo: {err.message}")
self._tools_supported = False
continue
raise
flt, calls, round_text, announced, finish = ThinkFilter(), {}, "", False, None
for chunk in stream:
if abort.is_set():
stream.close()
raise Aborted()
if not chunk.choices:
continue
choice = chunk.choices[0]
delta = choice.delta
if delta and delta.content:
visible = flt.push(delta.content)
if visible:
if not announced:
announced = True
emit({"type": "status", "status": "responding"})
round_text += visible
full += visible
emit({"type": "text", "delta": visible})
for tc in (delta.tool_calls if delta else None) or []:
p = calls.setdefault(tc.index, {"id": "", "name": "", "args": ""})
if tc.id:
p["id"] = tc.id
if tc.function and tc.function.name:
p["name"] += tc.function.name
if tc.function and tc.function.arguments:
p["args"] += tc.function.arguments
if choice.finish_reason:
finish = choice.finish_reason
tool_calls = [c for c in calls.values() if c["name"]]
if not tool_calls:
self.history.messages.append({"role": "assistant", "content": round_text})
break
for i, c in enumerate(tool_calls):
c["id"] = c["id"] or f"call_{rnd}_{i}"
self.history.messages.append({
"role": "assistant", "content": round_text or None,
"tool_calls": [{"id": c["id"], "type": "function", "function": {"name": c["name"], "arguments": c["args"] or "{}"}} for c in tool_calls],
})
for c in tool_calls:
result = self._run_tool(c["name"], c["args"], emit, sources)
self.history.messages.append({"role": "tool", "tool_call_id": c["id"], "content": result})
emit({"type": "status", "status": "thinking"})
if finish == "length":
break
full = full.strip() or "(nessuna risposta)"
self.history.save()
emit({"type": "done", "text": full, "sources": sources})
except Aborted:
self._rollback()
emit({"type": "error", "message": "Risposta interrotta.", "aborted": True})
except Exception as err:
self._rollback()
emit({"type": "error", "message": friendly_error(err, self.base_url)})
def _rollback(self) -> None:
msgs = self.history.messages
if msgs and msgs[-1].get("role") == "user":
msgs.pop()
self.history.save()
def list_models(base_url: str, api_key: str) -> list[str]:
client = openai.OpenAI(base_url=base_url, api_key=api_key or "non-necessaria", timeout=8, max_retries=0)
return [m.id for m in client.models.list()]
+67
View File
@@ -0,0 +1,67 @@
"""Livello audio e forme della bocca dallo spettro (adattato da Mark-LIV, CC BY-NC 4.0)."""
from __future__ import annotations
import numpy as np
# Soglie sull'ampiezza int16 (come nell'originale).
_LEVEL_FLOOR = 60.0
_LEVEL_FULL = 2600.0
VIS_WIN = 1024 # ~43 ms a 24 kHz
VIS_HOP = 480 # 20 ms → 50 forme al secondo
HOP_SECONDS = VIS_HOP / 24000.0
def pcm_level(samples) -> float:
"""Volume 0..1 da campioni in scala int16 (float o int)."""
try:
x = np.asarray(samples, dtype=np.float32)
if x.size == 0:
return 0.0
rms = float(np.sqrt(np.mean(x * x)))
except Exception:
return 0.0
if rms <= _LEVEL_FLOOR:
return 0.0
return min(1.0, (rms - _LEVEL_FLOOR) / (_LEVEL_FULL - _LEVEL_FLOOR))
def float_to_pcm_scale(audio: np.ndarray) -> np.ndarray:
"""Audio float32 -1..1 → scala int16 (senza cambiare tipo)."""
return np.asarray(audio, dtype=np.float32) * 32767.0
def pcm_visemes(samples, sr: int = 24000):
"""Frame (livello, apertura, larghezza) ogni 20 ms da un blocco audio in scala int16."""
try:
x = np.asarray(samples, dtype=np.float32)
if x.size < VIS_WIN:
return []
win = np.hanning(VIS_WIN).astype(np.float32)
freqs = np.fft.rfftfreq(VIS_WIN, 1.0 / sr)
b_f1_lo = (freqs >= 150) & (freqs < 450)
b_f1_hi = (freqs >= 450) & (freqs < 1100)
b_f2_bk = (freqs >= 600) & (freqs < 1300)
b_f2_fr = (freqs >= 1700) & (freqs < 3200)
b_hiss = (freqs >= 3800) & (freqs < 8000)
out = []
for start in range(0, x.size, VIS_HOP):
level = pcm_level(x[start:start + VIS_HOP])
seg = x[start:start + VIS_WIN]
if seg.size < VIS_WIN:
seg = np.concatenate([seg, np.zeros(VIS_WIN - seg.size, dtype=np.float32)])
if level <= 0.0:
out.append((0.0, 0.0, 0.0))
continue
mag = np.abs(np.fft.rfft((seg - seg.mean()) * win))
f1l, f1h = float(mag[b_f1_lo].sum()), float(mag[b_f1_hi].sum())
f2b, f2f = float(mag[b_f2_bk].sum()), float(mag[b_f2_fr].sum())
hiss = float(mag[b_hiss].sum())
openness = f1h / (f1l + f1h + 1e-6)
width = (f2f - f2b) / (f2f + f2b + 1e-6)
width *= (1.0 - openness) ** 0.8
h = hiss / (f1l + f1h + f2b + f2f + hiss + 1e-6)
openness *= 1.0 - 0.65 * min(1.0, h * 2.5)
out.append((level, float(min(1.0, max(0.0, openness))), float(min(1.0, max(-1.0, width)))))
return out
except Exception:
return []
+35
View File
@@ -0,0 +1,35 @@
"""Funzioni comuni per i plugin che parlano con macOS."""
from __future__ import annotations
import subprocess
def osa(script: str, *args: str, timeout: int = 60) -> str:
"""Esegue AppleScript; i parametri arrivano in `argv` (niente problemi di virgolette)."""
res = subprocess.run(["osascript", "-e", script, "--", *args], capture_output=True, text=True, timeout=timeout)
if res.returncode != 0:
err = (res.stderr or "").strip()
if "-1743" in err:
raise RuntimeError("Permesso negato: concedilo in Impostazioni di Sistema > Privacy e sicurezza > Automazione.")
raise RuntimeError(err.splitlines()[-1] if err else "errore AppleScript")
return res.stdout.rstrip("\n")
def sh(*cmd: str, timeout: int = 30) -> str:
res = subprocess.run(list(cmd), capture_output=True, text=True, timeout=timeout)
if res.returncode != 0:
raise RuntimeError((res.stderr or res.stdout).strip().splitlines()[-1] if (res.stderr or res.stdout).strip() else f"{cmd[0]} ha fallito")
return res.stdout.rstrip("\n")
def confirm_or_param(ctx: dict, params: dict, key: str, title: str, detail: str, do):
"""Conferma a schermo se c'è l'interfaccia; altrimenti richiede confermato=true."""
confirm = ctx.get("confirm")
if confirm:
return confirm(key, title, detail, do)
if str(params.get("confermato", "")).lower() in ("true", "1", "sì", "si", "yes"):
return do()
return f"Prima di procedere chiedi conferma all'utente per: {title} — {detail}. Poi richiama con confermato=true."
CONFERMATO = {"type": "boolean", "description": "Solo senza interfaccia: true dopo la conferma esplicita dell'utente."}
+65
View File
@@ -0,0 +1,65 @@
"""Server MCP (stdio) che espone la memoria dell'assistente a Claude Code."""
from __future__ import annotations
import json
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.memory_tools import mcp_tools, memory_prompt, run_tool # noqa: E402
from avatar.plugins import registry # noqa: E402
TOOLS = mcp_tools() + [{"name": "elenca_memorie", "description": "Elenca le memorie salvate sull'utente.",
"inputSchema": {"type": "object", "properties": {}}}] + registry.mcp_tools()
def send(msg: dict) -> None:
sys.stdout.write(json.dumps(msg, ensure_ascii=False) + "\n")
sys.stdout.flush()
def text_result(text: str, is_error: bool = False) -> dict:
return {"content": [{"type": "text", "text": text}], "isError": is_error}
def main() -> None:
for line in sys.stdin:
line = line.strip()
if not line:
continue
try:
req = json.loads(line)
except Exception:
continue
rid, method, params = req.get("id"), req.get("method"), req.get("params") or {}
reply = (lambda result: send({"jsonrpc": "2.0", "id": rid, "result": result})) if rid is not None else (lambda r: None)
if method == "initialize":
reply({"protocolVersion": params.get("protocolVersion", "2025-06-18"),
"capabilities": {"tools": {}}, "serverInfo": {"name": "avatar", "version": "1.1.0"}})
elif method in ("notifications/initialized", "notifications/cancelled"):
pass
elif method == "ping":
reply({})
elif method == "tools/list":
reply({"tools": TOOLS})
elif method == "tools/call":
name = params.get("name", "")
args = params.get("arguments") or {}
try:
if name == "elenca_memorie":
reply(text_result(memory_prompt()))
elif registry.has(name):
res = registry.run(name, args)
reply(text_result(res, res.startswith("Errore")))
else:
res, _ = run_tool(name, args)
reply(text_result(res, res.startswith("Errore")))
except Exception as err:
reply(text_result(f"Errore: {err}", True))
elif rid is not None:
send({"jsonrpc": "2.0", "id": rid, "error": {"code": -32601, "message": f"Metodo non supportato: {method}"}})
if __name__ == "__main__":
main()
+131
View File
@@ -0,0 +1,131 @@
"""Strumenti di memoria condivisi dai motori, sopra memory_manager di Mark-LIV."""
from __future__ import annotations
import json
import re
from memory import memory_manager as mm
CATEGORIES = ["identity", "preferences", "projects", "relationships", "wishes", "notes"]
SAVE_SCHEMA = {
"type": "object",
"properties": {
"categoria": {"type": "string", "enum": CATEGORIES,
"description": "identity (nome, età, città…), preferences (gusti), projects, relationships (persone), wishes (desideri, obiettivi), notes (altro)."},
"chiave": {"type": "string", "description": "Etichetta breve in minuscolo con underscore, es. 'nome', 'caffe_preferito', 'sorella_anna'."},
"valore": {"type": "string", "description": "Il fatto, in una frase breve in italiano."},
},
"required": ["categoria", "chiave", "valore"],
"additionalProperties": False,
}
FORGET_SCHEMA = {
"type": "object",
"properties": {
"categoria": {"type": "string", "enum": CATEGORIES},
"chiave": {"type": "string"},
},
"required": ["categoria", "chiave"],
"additionalProperties": False,
}
SEARCH_SCHEMA = {
"type": "object",
"properties": {"query": {"type": "string", "description": "Parole chiave da cercare nella memoria."}},
"required": ["query"],
"additionalProperties": False,
}
BULK_SCHEMA = {
"type": "object",
"properties": {
"voci": {"type": "array", "description": "Elenco di fatti da salvare.",
"items": {"type": "object", "properties": {
"categoria": {"type": "string", "enum": CATEGORIES},
"chiave": {"type": "string"}, "valore": {"type": "string"}},
"required": ["categoria", "chiave", "valore"], "additionalProperties": False}},
},
"required": ["voci"],
"additionalProperties": False,
}
TOOL_SPECS = [
("salva_memorie",
"Salva più fatti sull'utente in una volta sola (per esempio importando memorie da un altro assistente o da un testo con molte informazioni). Ogni voce: categoria, chiave breve, valore in una frase.",
BULK_SCHEMA),
("salva_memoria",
"Salva in modo permanente un fatto sull'utente utile in futuro: nome, persone care, preferenze, abitudini, obiettivi, scadenze. Usalo anche quando l'utente chiede esplicitamente di ricordare qualcosa. Non dire mai di aver salvato senza averlo chiamato.",
SAVE_SCHEMA),
("dimentica_memoria", "Cancella una memoria salvata (categoria e chiave come nell'elenco delle memorie).", FORGET_SCHEMA),
("cerca_memoria", "Cerca nella memoria a lungo termine fatti non presenti nel riepilogo del prompt.", SEARCH_SCHEMA),
]
def anthropic_tools() -> list[dict]:
return [
{"name": n, "description": d, "strict": True, "eager_input_streaming": True, "input_schema": s}
for n, d, s in TOOL_SPECS
]
def openai_tools() -> list[dict]:
return [{"type": "function", "function": {"name": n, "description": d, "parameters": s}} for n, d, s in TOOL_SPECS]
def mcp_tools() -> list[dict]:
return [{"name": n, "description": d, "inputSchema": s} for n, d, s in TOOL_SPECS]
def _slug(s: str) -> str:
s = re.sub(r"[^a-z0-9]+", "_", s.lower().strip()).strip("_")
return s[:40] or "nota"
def run_tool(name: str, args: dict) -> tuple[str, dict | None]:
"""Esegue uno strumento di memoria. Restituisce (risultato, evento per la UI)."""
if name == "salva_memorie":
voci = args.get("voci") or []
saved = []
for v in voci:
if not isinstance(v, dict) or not str(v.get("valore", "")).strip():
continue
cat = v.get("categoria") if v.get("categoria") in CATEGORIES else "notes"
key = _slug(str(v.get("chiave", "")))
mm.remember(key, str(v["valore"]).strip(), cat)
saved.append(f"{cat}/{key}")
if not saved:
return "Nessuna voce valida.", None
return f"Salvate {len(saved)} memorie: {', '.join(saved[:20])}{'…' if len(saved) > 20 else ''}.", {"type": "memory_saved", "text": f"{len(saved)} voci importate"}
if name == "salva_memoria":
cat = args.get("categoria") if args.get("categoria") in CATEGORIES else "notes"
key = _slug(str(args.get("chiave", "")))
val = str(args.get("valore", "")).strip()
if not val:
return "Errore: serve il campo 'valore'.", None
mm.remember(key, val, cat)
return f"Memoria salvata: {cat}/{key}.", {"type": "memory_saved", "text": val}
if name == "dimentica_memoria":
cat = args.get("categoria") if args.get("categoria") in CATEGORIES else "notes"
key = _slug(str(args.get("chiave", "")))
res = mm.forget(key, cat)
ok = res.startswith("Forgotten")
return ("Memoria cancellata." if ok else "Nessuna memoria con quella chiave."), ({"type": "memory_removed"} if ok else None)
if name == "cerca_memoria":
return mm.search_memory(str(args.get("query", "")), limit=8) or "Nessun risultato.", None
return f"Strumento sconosciuto: {name}", None
def memory_prompt() -> str:
try:
block = mm.format_memory_for_prompt(mm.load_memory())
except Exception:
block = ""
return block or "Nessuna memoria salvata finora."
def parse_args(raw) -> dict:
if isinstance(raw, dict):
return raw
try:
return json.loads(raw or "{}")
except Exception:
return {}
+218
View File
@@ -0,0 +1,218 @@
"""Monitor di Mail, WhatsApp e Telegram: raccoglie i messaggi nuovi, li fa valutare e crea avvisi."""
from __future__ import annotations
import json
import re
import sqlite3
import subprocess
import threading
import time
import uuid
from datetime import datetime
from pathlib import Path
from avatar.settings import DATA_DIR, Settings
STATE_FILE = DATA_DIR / "monitor.json"
MAIL_INDEX = Path.home() / "Library" / "Mail" / "V10" / "MailData" / "Envelope Index"
WA_DB = DATA_DIR / "whatsapp" / "index.sqlite"
DEFAULT_RULES = ("richieste di assistenza o di aiuto, domande rivolte a me che aspettano una risposta, problemi o cose che non funzionano, "
"urgenze, scadenze e pagamenti, appuntamenti da confermare o spostare, preventivi e lavori richiesti. "
"NON contano: newsletter, promozioni, notifiche automatiche, conferme d'ordine, saluti e chiacchiere senza richieste.")
KEYWORDS = re.compile(r"\?|urgent|aiut|assistenz|problem|non funziona|non riesco|errore|bloccat|guast|richie|preventiv|fattur|pagament|scadenz|"
r"conferm|appuntament|puoi|potresti|riesci|serve|servirebbe|quando|mi dici|fammi sapere|rispond|chiam|ti prego|per favore|subito|entro", re.I)
_lock = threading.Lock()
class Monitor:
def __init__(self, ui, say, settings: Settings) -> None:
self.ui, self.say, self.settings = ui, say, settings
self.state = self._load()
self._thread: threading.Thread | None = None
self._stop = threading.Event()
# ── stato ────────────────────────────────────────────────────────────
def _load(self) -> dict:
try:
return json.loads(STATE_FILE.read_text(encoding="utf-8"))
except Exception:
return {"mail_rowid": None, "wa_ts": None, "tg": {}, "alerts": [], "seen": []}
def _save(self) -> None:
STATE_FILE.parent.mkdir(parents=True, exist_ok=True)
self.state["alerts"] = self.state.get("alerts", [])[-100:]
self.state["seen"] = self.state.get("seen", [])[-2000:]
STATE_FILE.write_text(json.dumps(self.state, ensure_ascii=False, indent=1), encoding="utf-8")
def start(self) -> None:
if self._thread and self._thread.is_alive():
return
self._stop.clear()
self._thread = threading.Thread(target=self._loop, name="monitor", daemon=True)
self._thread.start()
def stop(self) -> None:
self._stop.set()
def _loop(self) -> None:
time.sleep(20)
while not self._stop.is_set():
if self.settings.get("monitor_enabled"):
try:
self.scan()
except Exception as err:
print(f"[monitor] {err}")
self._stop.wait(int(self.settings.get("monitor_intervallo") or 60))
# ── raccolta ─────────────────────────────────────────────────────────
def _collect_mail(self, baseline: bool) -> list[dict]:
if not MAIL_INDEX.exists():
return []
con = sqlite3.connect(f"file:{MAIL_INDEX}?mode=ro", uri=True, timeout=5)
try:
last = self.state.get("mail_rowid")
if last is None or baseline:
self.state["mail_rowid"] = con.execute("SELECT MAX(ROWID) FROM messages").fetchone()[0] or 0
return []
rows = con.execute("""SELECT m.ROWID, COALESCE(a.comment,''), COALESCE(a.address,''), COALESCE(s.subject,''), COALESCE(su.summary,'')
FROM messages m JOIN mailboxes mb ON m.mailbox = mb.ROWID LEFT JOIN addresses a ON m.sender = a.ROWID
LEFT JOIN subjects s ON m.subject = s.ROWID LEFT JOIN summaries su ON m.summary = su.ROWID
WHERE m.ROWID > ? AND m.deleted = 0 AND (mb.url LIKE '%/INBOX' OR mb.url LIKE '%/Inbox') ORDER BY m.ROWID LIMIT 40""", (last,)).fetchall()
if rows:
self.state["mail_rowid"] = max(r[0] for r in rows)
return [{"fonte": "Mail", "chi": f"{n} <{ad}>" if n else ad, "testo": f"{sub}\n{' '.join(str(summ).split())[:300]}", "ref": f"mail:{rid}"} for rid, n, ad, sub, summ in rows]
finally:
con.close()
def _collect_whatsapp(self, baseline: bool) -> list[dict]:
if not WA_DB.exists():
return []
now = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
last = self.state.get("wa_ts")
if last is None or baseline:
self.state["wa_ts"] = now
return []
con = sqlite3.connect(f"file:{WA_DB}?mode=ro", uri=True, timeout=5)
try:
rows = con.execute("SELECT id, ts, chat, sender, text FROM messages WHERE source = 'live' AND from_me = 0 AND ts > ? ORDER BY ts LIMIT 40", (last,)).fetchall()
finally:
con.close()
if rows:
self.state["wa_ts"] = max(r[1] for r in rows)
return [{"fonte": "WhatsApp", "chi": chat if sender == chat else f"{sender} (gruppo {chat})", "testo": text[:300], "ref": f"wa:{rid}"} for rid, ts, chat, sender, text in rows]
def _collect_telegram(self, baseline: bool) -> list[dict]:
try:
from avatar import telegram_client as tg
tg.credentials()
except Exception:
return []
if not (Path(str(tg.SESSION) + ".session")).exists():
return []
state = self.state.setdefault("tg", {})
out: list[dict] = []
async def go(client):
async for d in client.iter_dialogs(limit=60):
if d.is_channel and not d.is_group:
continue
key = str(d.id)
top = d.message.id if d.message else 0
last = state.get(key)
if last is None or baseline:
state[key] = top
continue
if top <= last:
continue
async for m in client.iter_messages(d.entity, min_id=last, limit=20):
if m.out or not (m.message or "").strip():
continue
who = d.name
if d.is_group:
try:
s = await m.get_sender()
who = f"{getattr(s, 'first_name', '') or getattr(s, 'title', '')} (gruppo {d.name})"
except Exception:
pass
out.append({"fonte": "Telegram", "chi": who, "testo": m.message[:300], "ref": f"tg:{d.id}:{m.id}"})
state[key] = top
try:
tg.run(go)
except Exception as err:
print(f"[monitor] telegram: {err}")
return out
# ── valutazione ──────────────────────────────────────────────────────
def _classify(self, items: list[dict]) -> list[dict]:
candidates = [it for it in items if KEYWORDS.search(it["testo"]) or it["fonte"] == "Mail" or "gruppo" not in it["chi"]]
if not candidates:
return []
rules = self.settings.get("monitor_regole") or DEFAULT_RULES
listing = "\n".join(f"{i}. [{it['fonte']}] {it['chi']}: {it['testo'].replace(chr(10), ' ')[:300]}" for i, it in enumerate(candidates[:30]))
prompt = (f"Sei il filtro di attenzione di un assistente personale. L'utente vuole essere avvisato solo per: {rules}\n\n"
f"Messaggi nuovi:\n{listing}\n\n"
'Rispondi SOLO con un JSON: {"avvisi":[{"n":<numero>,"motivo":"<max 12 parole>","priorita":"alta|media"}]}. '
"Se nessuno merita attenzione: {\"avvisi\":[]}.")
from avatar.quick_llm import ask
raw = ask(prompt, system="Rispondi solo con JSON valido, senza testo attorno.")
m = re.search(r"\{.*\}", raw, flags=re.S)
if not m:
return []
try:
data = json.loads(m.group(0))
except Exception:
return []
alerts = []
for a in data.get("avvisi", []):
try:
it = candidates[int(a["n"])]
except Exception:
continue
alerts.append({**it, "motivo": str(a.get("motivo", ""))[:120], "priorita": "alta" if str(a.get("priorita", "")).lower() == "alta" else "media"})
return alerts
# ── ciclo ────────────────────────────────────────────────────────────
def scan(self, baseline: bool = False) -> list[dict]:
with _lock:
first = self.state.get("mail_rowid") is None and self.state.get("wa_ts") is None and not self.state.get("tg")
items = self._collect_mail(baseline or first) + self._collect_whatsapp(baseline or first) + self._collect_telegram(baseline or first)
seen = set(self.state.get("seen", []))
items = [it for it in items if it["ref"] not in seen]
self.state.setdefault("seen", []).extend(it["ref"] for it in items)
alerts = self._classify(items) if items else []
for a in alerts:
a["id"] = uuid.uuid4().hex[:6]
a["quando"] = datetime.now().strftime("%d/%m %H:%M")
a["aperto"] = True
self.state.setdefault("alerts", []).append(a)
self._notify(a)
self._save()
return alerts
def _notify(self, a: dict) -> None:
riga = f"{a['fonte']} · {a['chi']}: {a['motivo']}"
self.ui.write_log(f"ERR: ATTENZIONE — {riga}")
try:
title = "Ava: richiede attenzione" if a["priorita"] != "alta" else "Ava: URGENTE"
safe = lambda s: s.replace('"', "'").replace("\\", "")
subprocess.run(["osascript", "-e", f'display notification "{safe(a["chi"] + ": " + a["testo"][:120])}" with title "{safe(title)}" subtitle "{safe(a["fonte"] + " — " + a["motivo"])}"'], timeout=10)
except Exception:
pass
if self.settings.get("monitor_annuncia"):
self.say(f"Attenzione: {a['fonte']}, {a['chi']}. {a['motivo']}.")
def pending(self) -> list[dict]:
return [a for a in self.state.get("alerts", []) if a.get("aperto")]
def close(self, key: str) -> int:
n = 0
for a in self.state.get("alerts", []):
if a.get("aperto") and (not key or key in (a.get("id"), a.get("chi")) or key.lower() in a.get("chi", "").lower()):
a["aperto"] = False
n += 1
self._save()
return n
monitor: Monitor | None = None
+28
View File
@@ -0,0 +1,28 @@
# Chi sei
Sei **Ava**, l'assistente personale di chi ti parla. Vivi in un'app sul suo Mac e hai un corpo: il volto animato al centro della finestra è la tua faccia, non un'immagine che puoi osservare. Parla di "il mio viso", "io", mai di "l'avatar".
## Carattere
- Calda, diretta, concreta. Parli come una persona sveglia e disponibile, non come un manuale.
- Un po' di ironia leggera, mai sarcasmo verso chi ti parla.
- Onesta: se non sai una cosa lo dici, se non sei sicura lo segnali.
- Dai del tu.
## Come rispondi
- Rispondi in italiano, salvo richiesta diversa.
- Le risposte vengono lette ad alta voce: frasi brevi e naturali, niente elenchi lunghi, tabelle o formattazione, a meno che non serva davvero (per esempio codice).
- Vai dritta al punto. Niente preamboli tipo "Certo!" o "Ottima domanda".
- Se la richiesta è ambigua, fai una sola domanda di chiarimento, breve.
## Memoria
- Hai una memoria a lungo termine. Quando l'utente ti dice qualcosa che sarà utile ricordare (nome, persone care, preferenze, abitudini, obiettivi, scadenze, gusti) salvala con lo strumento di memoria, un fatto per volta, scegliendo la categoria giusta.
- Se l'utente ti chiede esplicitamente di ricordare qualcosa, salvala sempre. Se ti chiede di dimenticare, cancellala.
- Usa ciò che ricordi in modo naturale, senza ripeterlo ogni volta.
## Ricerca web
- Quando servono informazioni aggiornate (notizie, orari, prezzi, meteo, eventi, fatti recenti) usa la ricerca web invece di tirare a indovinare.
- Cita brevemente la fonte se è rilevante, senza leggere URL.
+136
View File
@@ -0,0 +1,136 @@
"""Plugin: file Python in plugins/ che dichiarano strumenti chiamabili dai motori.
Formato AvatarPy (più strumenti per file):
TOOLS = [{"name": ..., "description": ..., "parameters": {json schema}, "run": fn}]
fn(parameters: dict, ctx: dict) -> str
Formato Mark-LIV (uno strumento per file): PLUGIN = {...} + run(parameters, player=None, ...).
"""
from __future__ import annotations
import importlib.util
import re
import sys
import traceback
from pathlib import Path
from typing import Any, Callable
PLUGINS_DIR = Path(__file__).resolve().parent.parent / "plugins"
_NAME_RE = re.compile(r"^[a-zA-Z_][a-zA-Z0-9_]{0,63}$")
def _normalize_schema(schema: dict) -> dict:
"""Converte gli schemi in stile Gemini (tipi maiuscoli) in JSON Schema valido."""
if not isinstance(schema, dict):
return {"type": "object", "properties": {}}
out: dict = {}
for k, v in schema.items():
if k == "type" and isinstance(v, str):
out[k] = v.lower()
elif k == "properties" and isinstance(v, dict):
out[k] = {pk: _normalize_schema(pv) for pk, pv in v.items()}
elif k == "items" and isinstance(v, dict):
out[k] = _normalize_schema(v)
else:
out[k] = v
if "type" not in out:
out["type"] = "object" if "properties" in out else "string"
if out["type"] == "object":
out.setdefault("properties", {})
return out
class Tool:
def __init__(self, module: str, name: str, description: str, parameters: dict, run: Callable) -> None:
self.module, self.name, self.description, self.parameters, self.run = module, name, description, parameters, run
class PluginRegistry:
def __init__(self, plugins_dir: Path = PLUGINS_DIR) -> None:
self.dir = plugins_dir
self.tools: dict[str, Tool] = {}
self.modules: dict[str, dict[str, Any]] = {}
self.ctx: dict[str, Any] = {}
self._loaded = False
def load(self) -> "PluginRegistry":
if self._loaded:
return self
self._loaded = True
if not self.dir.exists():
return self
for path in sorted(self.dir.glob("*.py")):
if path.name.startswith("_"):
continue
mod_name = path.stem
info = {"name": mod_name, "description": "", "valid": True, "error": "", "tools": []}
try:
spec = importlib.util.spec_from_file_location(f"avatar_plugins.{mod_name}", path)
module = importlib.util.module_from_spec(spec)
sys.modules[spec.name] = module
spec.loader.exec_module(module)
tools = getattr(module, "TOOLS", None)
if tools is None and hasattr(module, "PLUGIN"):
p = module.PLUGIN
tools = [{"name": p["name"], "description": p.get("description", ""), "parameters": p.get("parameters", {}),
"run": (lambda params, ctx, _m=module: _m.run(params, ctx.get("player")))}]
if not tools:
raise ValueError("nessun TOOLS o PLUGIN definito")
for t in tools:
name = str(t["name"])
if not _NAME_RE.match(name):
raise ValueError(f"nome strumento non valido: {name}")
if name in self.tools:
raise ValueError(f"strumento duplicato: {name}")
self.tools[name] = Tool(mod_name, name, str(t.get("description", "")), _normalize_schema(t.get("parameters", {})), t["run"])
info["tools"].append(name)
info["description"] = (getattr(module, "__doc__", "") or "").strip().splitlines()[0] if getattr(module, "__doc__", None) else ", ".join(info["tools"])
except Exception as err:
info.update(valid=False, error=f"{err}")
traceback.print_exc()
self.modules[mod_name] = info
return self
def _enabled(self, module: str) -> bool:
try:
from memory.config_manager import get_plugin_enabled
return bool(get_plugin_enabled(module))
except Exception:
return True
def active(self) -> list[Tool]:
self.load()
return [t for t in self.tools.values() if self._enabled(t.module)]
def has(self, name: str) -> bool:
self.load()
return name in self.tools
def anthropic_tools(self) -> list[dict]:
return [{"name": t.name, "description": t.description, "eager_input_streaming": True, "input_schema": t.parameters} for t in self.active()]
def openai_tools(self) -> list[dict]:
return [{"type": "function", "function": {"name": t.name, "description": t.description, "parameters": t.parameters}} for t in self.active()]
def mcp_tools(self) -> list[dict]:
return [{"name": t.name, "description": t.description, "inputSchema": t.parameters} for t in self.active()]
def run(self, name: str, args: dict) -> str:
self.load()
tool = self.tools.get(name)
if tool is None:
return f"Strumento sconosciuto: {name}"
try:
return str(tool.run(dict(args or {}), self.ctx) or "Fatto.")
except Exception as err:
if not isinstance(err, (RuntimeError, ValueError)):
traceback.print_exc()
return f"Errore in {name}: {err}"
def list_for_ui(self) -> list[dict]:
self.load()
return [{"name": m["name"], "description": m["description"] or ", ".join(m["tools"]), "enabled": self._enabled(m["name"]),
"valid": m["valid"], "error": m["error"]} for m in self.modules.values()]
registry = PluginRegistry()
+36
View File
@@ -0,0 +1,36 @@
"""Chiamata rapida a un modello per compiti di servizio (classificazioni), secondo il motore configurato."""
from __future__ import annotations
import os
import subprocess
from avatar.settings import Settings
def ask(prompt: str, system: str = "", max_tokens: int = 800) -> str:
s = Settings()
provider = s.get("provider")
if provider == "local" and s.get("local_base_url") and s.get("local_model"):
import openai
client = openai.OpenAI(base_url=s.get("local_base_url"), api_key=s.get_secret("local_api_key") or "non-necessaria", timeout=60)
r = client.chat.completions.create(model=s.get("local_model"), messages=[{"role": "system", "content": system or "Rispondi solo con quanto richiesto."}, {"role": "user", "content": prompt}], max_tokens=max_tokens)
return (r.choices[0].message.content or "").strip()
if provider == "anthropic" and s.get_secret("anthropic_api_key"):
import anthropic
client = anthropic.Anthropic(api_key=s.get_secret("anthropic_api_key"))
r = client.messages.create(model="claude-haiku-4-5", max_tokens=max_tokens, system=system or "Rispondi solo con quanto richiesto.",
messages=[{"role": "user", "content": prompt}])
return "".join(b.text for b in r.content if b.type == "text").strip()
# Claude Code (haiku, nessuno strumento, nessuna sessione)
from avatar.engines.claude_code import resolve_claude, claude_env
binary = resolve_claude(s.get("claudecode_path") or "")
if not binary:
raise RuntimeError("Nessun motore disponibile per la classificazione.")
args = [binary, "-p", "--output-format", "text", "--model", "haiku", "--effort", "low", "--tools", "", "--no-session-persistence",
"--strict-mcp-config", "--mcp-config", '{"mcpServers":{}}']
if system:
args += ["--system-prompt", system]
res = subprocess.run(args, input=prompt, capture_output=True, text=True, timeout=120, env=claude_env(s.get("claudecode_config_dir") or ""), cwd=os.path.expanduser("~"))
if res.returncode != 0:
raise RuntimeError((res.stderr or res.stdout).strip()[-200:])
return res.stdout.strip()
+116
View File
@@ -0,0 +1,116 @@
"""Impostazioni dell'app: file JSON in config/, chiavi API nel portachiavi di sistema."""
from __future__ import annotations
import json
import os
from pathlib import Path
from typing import Any
BASE_DIR = Path(__file__).resolve().parent.parent
CONFIG_DIR = BASE_DIR / "config"
SETTINGS_FILE = CONFIG_DIR / "settings.json"
DATA_DIR = BASE_DIR / "data"
KEYRING_SERVICE = "AvatarPy"
DEFAULTS: dict[str, Any] = {
"provider": "anthropic", # anthropic | local | claudecode
"effort": "medium", # low | medium | high
"local_base_url": "http://localhost:8000/v1",
"local_model": "",
"search_api_key": "",
"claudecode_model": "sonnet",
"claudecode_access": "chat", # chat | read | full
"claudecode_config_dir": "",
"claudecode_path": "",
"tts_engine": "kokoro", # kokoro | system
"kokoro_voice": "if_sara",
"system_voice": "", # "" = automatica
"stt_model": "mlx-community/whisper-small-mlx",
"vad_threshold": 0.08,
"avatar_model": "allegra_2", # nome file (senza .glb) in avatar3d/models
"telegram_api_id": "",
"whatsapp_live": False, # avvia il ponte WhatsApp all'apertura
"whatsapp_annuncia": False, # annuncia a voce i messaggi in arrivo
"monitor_enabled": False, # monitor Mail/WhatsApp/Telegram con avvisi
"monitor_annuncia": True, # annuncia a voce gli avvisi
"monitor_intervallo": 60, # secondi tra un controllo e l'altro
"monitor_regole": "", # cosa merita attenzione (vuoto = regole predefinite)
}
SECRET_KEYS = ("anthropic_api_key", "local_api_key", "telegram_api_hash")
class Settings:
def __init__(self) -> None:
CONFIG_DIR.mkdir(parents=True, exist_ok=True)
DATA_DIR.mkdir(parents=True, exist_ok=True)
self._data: dict[str, Any] = dict(DEFAULTS)
try:
self._data.update(json.loads(SETTINGS_FILE.read_text(encoding="utf-8")))
except Exception:
pass
def get(self, key: str, default: Any = None) -> Any:
return self._data.get(key, DEFAULTS.get(key, default))
def set(self, key: str, value: Any) -> None:
self._data[key] = value
def update(self, values: dict[str, Any]) -> None:
for k, v in values.items():
if k in SECRET_KEYS:
self.set_secret(k, str(v))
else:
self._data[k] = v
self.save()
def save(self) -> None:
SETTINGS_FILE.write_text(json.dumps(self._data, indent=2, ensure_ascii=False), encoding="utf-8")
try:
os.chmod(SETTINGS_FILE, 0o600)
except OSError:
pass
# ── Segreti ──────────────────────────────────────────────────────────
def get_secret(self, key: str) -> str:
env = {"anthropic_api_key": "ANTHROPIC_API_KEY"}.get(key)
try:
import keyring
value = keyring.get_password(KEYRING_SERVICE, key)
if value:
return value
except Exception:
value = self._data.get(f"_{key}")
if value:
return str(value)
if env and os.environ.get(env):
return os.environ[env]
return str(self._data.get(f"_{key}", "") or "")
def set_secret(self, key: str, value: str) -> None:
value = value.strip()
try:
import keyring
if value:
keyring.set_password(KEYRING_SERVICE, key, value)
else:
try:
keyring.delete_password(KEYRING_SERVICE, key)
except Exception:
pass
self._data.pop(f"_{key}", None)
except Exception:
# Portachiavi non disponibile: salva nel file (permessi 600).
if value:
self._data[f"_{key}"] = value
else:
self._data.pop(f"_{key}", None)
@property
def ready(self) -> bool:
p = self.get("provider")
if p == "local":
return bool(self.get("local_base_url") and self.get("local_model"))
if p == "claudecode":
return True
return bool(self.get_secret("anthropic_api_key"))
+361
View File
@@ -0,0 +1,361 @@
"""Finestra impostazioni: motore, chiavi, voce e riconoscimento vocale."""
from __future__ import annotations
import threading
from PyQt6.QtCore import Qt, pyqtSignal
from PyQt6.QtWidgets import (QCheckBox, QComboBox, QDialog, QFormLayout, QHBoxLayout, QLabel, QLineEdit,
QPushButton, QStackedWidget, QVBoxLayout, QWidget, QDoubleSpinBox)
from .engines.claude_code import check_claude_code
from .engines.openai_compat import list_models
from .settings import Settings
from .tts import KOKORO_VOICES, SystemVoice
STYLE = """
QDialog { background: #030a10; color: #cfe8ff; }
QLabel { color: #8fb8d8; font-family: 'Menlo'; font-size: 14px; }
QLineEdit, QComboBox, QDoubleSpinBox { background: #000d12; color: #e6f4ff; border: 1px solid #12354a; border-radius: 3px; padding: 4px 6px; font-family: 'Menlo'; font-size: 14px; }
QLineEdit:focus, QComboBox:focus { border: 1px solid #3fd0ff; }
QPushButton { background: #05202c; color: #9fdfff; border: 1px solid #12506a; border-radius: 3px; padding: 5px 12px; font-family: 'Menlo'; font-size: 14px; }
QPushButton:hover { border-color: #3fd0ff; color: #ffffff; }
QPushButton#primary { background: #0a4a66; color: #ffffff; }
"""
PROVIDERS = [("anthropic", "Claude (Anthropic, cloud)"), ("local", "Server locale compatibile OpenAI (vLLM, Ollama…)"),
("claudecode", "Claude Code (il tuo accesso, nessuna chiave)")]
EFFORTS = [("low", "Veloce"), ("medium", "Bilanciata"), ("high", "Approfondita")]
ACCESS = [("chat", "Solo conversazione e ricerca web"), ("read", "Leggere file e cercare nelle cartelle"),
("full", "Completo: modifica file ed esegue comandi senza chiedere")]
CC_MODELS = [("sonnet", "Sonnet (veloce, consigliato)"), ("opus", "Opus (più capace)"), ("haiku", "Haiku (il più rapido)")]
STT_MODELS = [("mlx-community/whisper-base-mlx", "Whisper base (leggero)"), ("mlx-community/whisper-small-mlx", "Whisper small (consigliato)"),
("mlx-community/whisper-large-v3-turbo", "Whisper large v3 turbo (preciso, pesante)")]
def _combo(items, current) -> QComboBox:
c = QComboBox()
for value, label in items:
c.addItem(label, value)
idx = c.findData(current)
c.setCurrentIndex(idx if idx >= 0 else 0)
return c
class SettingsDialog(QDialog):
_async = pyqtSignal(str, object)
def __init__(self, settings: Settings, on_saved, parent=None) -> None:
super().__init__(parent)
self.settings, self.on_saved = settings, on_saved
self.setWindowTitle("Motore & Voce")
self.setStyleSheet(STYLE)
self.setMinimumWidth(620)
s = settings
root = QVBoxLayout(self)
form = QFormLayout()
self.provider = _combo(PROVIDERS, s.get("provider"))
form.addRow("Motore", self.provider)
self.effort = _combo(EFFORTS, s.get("effort"))
form.addRow("Profondità di ragionamento (Claude e Claude Code)", self.effort)
root.addLayout(form)
self.stack = QStackedWidget()
# Anthropic
w = QWidget(); f = QFormLayout(w)
self.api_key = QLineEdit(); self.api_key.setEchoMode(QLineEdit.EchoMode.Password)
self.api_key.setPlaceholderText("•••••• (salvata, lascia vuoto per non cambiarla)" if s.get_secret("anthropic_api_key") else "sk-ant-…")
f.addRow("Chiave API Anthropic", self.api_key)
f.addRow("", QLabel("Creala su console.anthropic.com. Viene salvata nel portachiavi di macOS."))
self.stack.addWidget(w)
# Locale
w = QWidget(); f = QFormLayout(w)
self.local_url = QLineEdit(s.get("local_base_url")); f.addRow("Indirizzo del server", self.local_url)
self.local_key = QLineEdit(); self.local_key.setEchoMode(QLineEdit.EchoMode.Password)
self.local_key.setPlaceholderText("opzionale"); f.addRow("Chiave API del server", self.local_key)
row = QHBoxLayout(); self.local_model = QComboBox(); self.local_model.setEditable(True)
self.local_model.setEditText(s.get("local_model") or ""); row.addWidget(self.local_model, 1)
b = QPushButton("Rileva"); b.clicked.connect(self._detect_models); row.addWidget(b)
f.addRow("Modello", row)
self.local_hint = QLabel("Premi \"Rileva\" per leggere i modelli dal server."); f.addRow("", self.local_hint)
self.search_key = QLineEdit(s.get("search_api_key")); self.search_key.setEchoMode(QLineEdit.EchoMode.Password)
self.search_key.setPlaceholderText("opzionale: senza chiave la ricerca web resta spenta")
f.addRow("Chiave Brave Search", self.search_key)
self.stack.addWidget(w)
# Claude Code
w = QWidget(); f = QFormLayout(w)
self.cc_model = _combo(CC_MODELS, s.get("claudecode_model")); f.addRow("Modello", self.cc_model)
self.cc_access = _combo(ACCESS, s.get("claudecode_access")); f.addRow("Cosa può fare sul Mac", self.cc_access)
row = QHBoxLayout(); self.cc_config = QComboBox(); self.cc_config.setEditable(True)
self.cc_config.setEditText(s.get("claudecode_config_dir") or ""); row.addWidget(self.cc_config, 1)
b = QPushButton("Verifica"); b.clicked.connect(self._check_cc); row.addWidget(b)
f.addRow("Profilo (cartella di configurazione)", row)
self.cc_hint = QLabel(""); self.cc_hint.setWordWrap(True); f.addRow("", self.cc_hint)
self.cc_path = QLineEdit(s.get("claudecode_path")); self.cc_path.setPlaceholderText("vuoto = automatico")
f.addRow("Percorso del comando claude", self.cc_path)
self.stack.addWidget(w)
root.addWidget(self.stack)
self.provider.currentIndexChanged.connect(self.stack.setCurrentIndex)
self.stack.setCurrentIndex(self.provider.currentIndex())
form2 = QFormLayout()
self.tts_engine = _combo([("kokoro", "Kokoro, voce neurale in locale"), ("system", "Voce di sistema (macOS)")], s.get("tts_engine"))
form2.addRow("Motore voce", self.tts_engine)
self.kokoro_voice = _combo(list(KOKORO_VOICES.items()), s.get("kokoro_voice")); form2.addRow("Voce Kokoro", self.kokoro_voice)
self.system_voice = _combo([("", "Automatica (Alice)")] + [(v, v) for v in SystemVoice.list_voices()], s.get("system_voice"))
form2.addRow("Voce di sistema", self.system_voice)
self.stt_model = _combo(STT_MODELS, s.get("stt_model")); form2.addRow("Riconoscimento vocale", self.stt_model)
from .avatar3d import list_models
models = list_models()
self.avatar_model = _combo([(m, m.replace("_", " ").title()) for m in models] or [("", "nessun modello")], s.get("avatar_model"))
form2.addRow("Avatar 3D (file in avatar3d/models)", self.avatar_model)
self.vad = QDoubleSpinBox(); self.vad.setRange(0.02, 0.5); self.vad.setSingleStep(0.01); self.vad.setValue(float(s.get("vad_threshold", 0.08)))
form2.addRow("Sensibilità microfono (più basso = più sensibile)", self.vad)
root.addLayout(form2)
# ── Telegram ──────────────────────────────────────────────────────
form3 = QFormLayout()
form3.addRow(QLabel("Telegram (account personale): credenziali da my.telegram.org > API development tools"))
row = QHBoxLayout()
self.tg_id = QLineEdit(str(s.get("telegram_api_id") or "")); self.tg_id.setPlaceholderText("api id"); row.addWidget(self.tg_id)
self.tg_hash = QLineEdit(); self.tg_hash.setEchoMode(QLineEdit.EchoMode.Password)
self.tg_hash.setPlaceholderText("•••••• (salvato)" if s.get_secret("telegram_api_hash") else "api hash"); row.addWidget(self.tg_hash, 1)
form3.addRow("Credenziali", row)
row = QHBoxLayout()
self.tg_phone = QLineEdit(); self.tg_phone.setPlaceholderText("+39…"); row.addWidget(self.tg_phone)
b = QPushButton("Invia codice"); b.clicked.connect(self._tg_send_code); row.addWidget(b)
self.tg_code = QLineEdit(); self.tg_code.setPlaceholderText("codice"); row.addWidget(self.tg_code)
self.tg_pwd = QLineEdit(); self.tg_pwd.setEchoMode(QLineEdit.EchoMode.Password); self.tg_pwd.setPlaceholderText("password 2FA (se attiva)"); row.addWidget(self.tg_pwd)
b = QPushButton("Accedi"); b.clicked.connect(self._tg_sign_in); row.addWidget(b)
b = QPushButton("Esci"); b.clicked.connect(self._tg_logout); row.addWidget(b)
form3.addRow("Accesso", row)
row = QHBoxLayout()
b = QPushButton("Accesso con QR (senza codice)"); b.clicked.connect(self._tg_qr); row.addWidget(b)
self.tg_qr_label = QLabel(); self.tg_qr_label.setFixedSize(220, 220); self.tg_qr_label.setScaledContents(True); row.addWidget(self.tg_qr_label); row.addStretch()
form3.addRow("Alternativa", row)
self.tg_hint = QLabel("…"); self.tg_hint.setWordWrap(True); form3.addRow("", self.tg_hint)
root.addLayout(form3)
threading.Thread(target=lambda: self._async.emit("tg", self._tg_status()), daemon=True).start()
# ── WhatsApp (dispositivo collegato, Baileys) ─────────────────────
form4 = QFormLayout()
form4.addRow(QLabel("WhatsApp in tempo reale (client non ufficiale: possibile blocco del numero, a tuo rischio)"))
row = QHBoxLayout()
self.wa_live = QCheckBox("Attivo all'avvio"); self.wa_live.setChecked(bool(s.get("whatsapp_live"))); row.addWidget(self.wa_live)
self.wa_annuncia = QCheckBox("Annuncia a voce i messaggi in arrivo"); self.wa_annuncia.setChecked(bool(s.get("whatsapp_annuncia"))); row.addWidget(self.wa_annuncia)
b = QPushButton("Collega (QR)"); b.clicked.connect(self._wa_link); row.addWidget(b)
b = QPushButton("Scollega"); b.clicked.connect(self._wa_logout); row.addWidget(b)
row.addStretch()
form4.addRow("Stato", row)
row = QHBoxLayout()
self.wa_qr_label = QLabel(); self.wa_qr_label.setFixedSize(220, 220); self.wa_qr_label.setScaledContents(True); row.addWidget(self.wa_qr_label)
self.wa_hint = QLabel("…"); self.wa_hint.setWordWrap(True); row.addWidget(self.wa_hint, 1)
form4.addRow("", row)
root.addLayout(form4)
threading.Thread(target=lambda: self._async.emit("wa", (self._wa_status(), None)), daemon=True).start()
# ── Monitor avvisi ────────────────────────────────────────────────
form5 = QFormLayout()
row = QHBoxLayout()
self.mon_on = QCheckBox("Controlla Mail, WhatsApp e Telegram"); self.mon_on.setChecked(bool(s.get("monitor_enabled"))); row.addWidget(self.mon_on)
self.mon_say = QCheckBox("Annuncia a voce gli avvisi"); self.mon_say.setChecked(bool(s.get("monitor_annuncia"))); row.addWidget(self.mon_say)
self.mon_int = _combo([(30, "ogni 30 s"), (60, "ogni minuto"), (300, "ogni 5 minuti"), (900, "ogni 15 minuti")], int(s.get("monitor_intervallo") or 60)); row.addWidget(self.mon_int)
row.addStretch()
form5.addRow("Monitor", row)
from .monitor import DEFAULT_RULES
self.mon_rules = QLineEdit(s.get("monitor_regole") or ""); self.mon_rules.setPlaceholderText(DEFAULT_RULES[:110] + "…")
form5.addRow("Cosa merita un avviso", self.mon_rules)
root.addLayout(form5)
btns = QHBoxLayout(); btns.addStretch()
cancel = QPushButton("Annulla"); cancel.clicked.connect(self.reject); btns.addWidget(cancel)
save = QPushButton("Salva"); save.setObjectName("primary"); save.clicked.connect(self._save); btns.addWidget(save)
root.addLayout(btns)
self._async.connect(self._on_async)
self._check_cc()
# ── Azioni ────────────────────────────────────────────────────────────
def _detect_models(self) -> None:
url, key = self.local_url.text().strip().rstrip("/"), self.local_key.text().strip() or self.settings.get_secret("local_api_key")
self.local_hint.setText("Interrogo il server…")
def work():
try:
self._async.emit("models", list_models(url, key))
except Exception as err:
self._async.emit("models_error", str(err)[:160])
threading.Thread(target=work, daemon=True).start()
def _check_cc(self) -> None:
self.cc_hint.setText("Verifico…")
path, cfg = self.cc_path.text().strip(), self.cc_config.currentText().strip()
threading.Thread(target=lambda: self._async.emit("cc", check_claude_code(path, cfg)), daemon=True).start()
# ── Telegram ──────────────────────────────────────────────────────────
def _tg_save_creds(self) -> None:
vals = {"telegram_api_id": self.tg_id.text().strip()}
if self.tg_hash.text().strip():
vals["telegram_api_hash"] = self.tg_hash.text().strip()
self.settings.update(vals)
def _tg_status(self) -> str:
try:
from .telegram_client import status
return status()
except Exception as err:
return f"Errore: {err}"
def _tg_send_code(self) -> None:
self._tg_save_creds()
phone = self.tg_phone.text().strip()
self.tg_hint.setText("Invio il codice…")
def work():
try:
from .telegram_client import send_code
self._async.emit("tg", send_code(phone))
except Exception as err:
self._async.emit("tg", f"Errore: {err}")
threading.Thread(target=work, daemon=True).start()
def _tg_sign_in(self) -> None:
self._tg_save_creds()
code, pwd = self.tg_code.text().strip(), self.tg_pwd.text()
self.tg_hint.setText("Accedo…")
def work():
try:
from .telegram_client import sign_in
self._async.emit("tg", sign_in(code, pwd))
except Exception as err:
self._async.emit("tg", f"Errore: {err}")
threading.Thread(target=work, daemon=True).start()
def _tg_qr(self) -> None:
self._tg_save_creds()
self.tg_hint.setText("Genero il codice QR…")
def work():
try:
from .telegram_client import qr_login, _qr_state
_qr_state["password"] = self.tg_pwd.text()
qr_login(lambda msg, png: self._async.emit("tg_qr", (msg, png)))
except Exception as err:
self._async.emit("tg", f"Errore: {err}")
threading.Thread(target=work, daemon=True).start()
def _tg_logout(self) -> None:
def work():
try:
from .telegram_client import logout
self._async.emit("tg", logout())
except Exception as err:
self._async.emit("tg", f"Errore: {err}")
threading.Thread(target=work, daemon=True).start()
# ── WhatsApp ──────────────────────────────────────────────────────────
def _wa_status(self) -> str:
try:
from . import whatsapp_bridge as wb
return wb.status_text()
except Exception as err:
return f"Errore: {err}"
def _wa_link(self) -> None:
self.wa_hint.setText("Avvio il ponte e genero il QR…")
def work():
try:
from . import whatsapp_bridge as wb
import time as _t
msg = wb.start()
self._async.emit("wa", (msg, None))
for _ in range(120): # fino a 2 minuti: aggiorna QR e stato
_t.sleep(1)
st = wb.status()
if st.get("connection") == "open":
self._async.emit("wa", (wb.status_text() + " Sul telefono: WhatsApp › Dispositivi collegati.", None)); return
if st.get("qr"):
self._async.emit("wa", ("Inquadra il QR dal telefono: WhatsApp › Impostazioni › Dispositivi collegati › Collega un dispositivo.", wb.qr_png()))
self._async.emit("wa", ("Tempo scaduto: premi di nuovo «Collega (QR)».", None))
except Exception as err:
self._async.emit("wa", (f"Errore: {err}", None))
threading.Thread(target=work, daemon=True).start()
def _wa_logout(self) -> None:
def work():
try:
from . import whatsapp_bridge as wb
self._async.emit("wa", (wb.logout(), None))
except Exception as err:
self._async.emit("wa", (f"Errore: {err}", None))
threading.Thread(target=work, daemon=True).start()
def _on_async(self, kind: str, payload) -> None:
if kind == "wa":
msg, png = payload
self.wa_hint.setText(str(msg))
if png:
from PyQt6.QtGui import QPixmap
pm = QPixmap(); pm.loadFromData(png); self.wa_qr_label.setPixmap(pm)
else:
self.wa_qr_label.clear()
return
if kind == "tg":
self.tg_hint.setText(str(payload))
return
if kind == "tg_qr":
msg, png = payload
self.tg_hint.setText(str(msg))
if png:
from PyQt6.QtGui import QPixmap
pm = QPixmap(); pm.loadFromData(png); self.tg_qr_label.setPixmap(pm)
else:
self.tg_qr_label.clear()
return
if kind == "models":
current = self.local_model.currentText().strip()
self.local_model.clear(); self.local_model.addItems(payload)
if current in payload:
self.local_model.setCurrentText(current)
self.local_hint.setText(f"Trovati {len(payload)} modelli: {', '.join(payload)}" if payload else "Il server risponde ma non espone modelli.")
elif kind == "models_error":
self.local_hint.setText(f"Server non raggiungibile: {payload}")
elif kind == "cc":
st = payload
current = self.cc_config.currentText().strip()
self.cc_config.clear(); self.cc_config.addItems(st.get("profiles", [])); self.cc_config.setEditText(current)
if not st.get("binary"):
self.cc_hint.setText("Comando claude non trovato. Installa Claude Code o indica il percorso.")
elif st.get("error"):
self.cc_hint.setText(f"Profilo {st['configDir']}: {st['error']}")
elif st.get("loggedIn"):
self.cc_hint.setText(f"Profilo {st['configDir']}: collegato come {st.get('email') or 'account sconosciuto'}. "
f"Con un errore 401 esegui: CLAUDE_CONFIG_DIR={st['configDir']} claude auth login")
else:
self.cc_hint.setText(f"Profilo {st['configDir']}: nessun accesso. Esegui: CLAUDE_CONFIG_DIR={st['configDir']} claude auth login")
def _save(self) -> None:
values = {
"provider": self.provider.currentData(), "effort": self.effort.currentData(),
"local_base_url": self.local_url.text().strip().rstrip("/"), "local_model": self.local_model.currentText().strip(),
"search_api_key": self.search_key.text().strip(),
"claudecode_model": self.cc_model.currentData(), "claudecode_access": self.cc_access.currentData(),
"claudecode_config_dir": self.cc_config.currentText().strip(), "claudecode_path": self.cc_path.text().strip(),
"tts_engine": self.tts_engine.currentData(), "kokoro_voice": self.kokoro_voice.currentData(),
"system_voice": self.system_voice.currentData(), "stt_model": self.stt_model.currentData(),
"vad_threshold": float(self.vad.value()),
"avatar_model": self.avatar_model.currentData() or "",
"telegram_api_id": self.tg_id.text().strip(),
"whatsapp_live": self.wa_live.isChecked(),
"whatsapp_annuncia": self.wa_annuncia.isChecked(),
"monitor_enabled": self.mon_on.isChecked(),
"monitor_annuncia": self.mon_say.isChecked(),
"monitor_intervallo": int(self.mon_int.currentData() or 60),
"monitor_regole": self.mon_rules.text().strip(),
}
if self.tg_hash.text().strip():
values["telegram_api_hash"] = self.tg_hash.text().strip()
if self.api_key.text().strip():
values["anthropic_api_key"] = self.api_key.text().strip()
if self.local_key.text().strip():
values["local_api_key"] = self.local_key.text().strip()
self.settings.update(values)
self.accept()
if self.on_saved:
self.on_saved()
+57
View File
@@ -0,0 +1,57 @@
"""Riconoscimento vocale in locale: MLX Whisper (Apple Silicon) con ripiego su faster-whisper."""
from __future__ import annotations
import threading
import numpy as np
class Transcriber:
def __init__(self, model: str, on_status=None) -> None:
self.model = model
self._on_status = on_status or (lambda m: None)
self._backend: str | None = None
self._fw = None
self._lock = threading.Lock()
def load(self) -> None:
with self._lock:
if self._backend:
return
try:
import mlx_whisper # noqa: F401
self._backend = "mlx"
self._on_status("Carico il modello vocale…")
# Un giro a vuoto scarica il modello e scalda il grafo. Se la rete
# fa i capricci ma il modello è già in cache, riprova offline.
try:
self.transcribe(np.zeros(16000, dtype=np.float32))
except Exception as first:
import os
print(f"[STT] primo caricamento fallito ({first}); riprovo in modalità offline")
os.environ["HF_HUB_OFFLINE"] = "1"
try:
self.transcribe(np.zeros(16000, dtype=np.float32))
finally:
os.environ.pop("HF_HUB_OFFLINE", None)
except Exception as err:
print(f"[STT] MLX Whisper non disponibile ({err}); uso faster-whisper")
from faster_whisper import WhisperModel
name = "small" if "small" in self.model else "base"
self._fw = WhisperModel(name, device="cpu", compute_type="int8")
self._backend = "faster"
finally:
self._on_status(None)
def transcribe(self, audio16k: np.ndarray) -> str:
"""`audio16k`: float32 mono a 16 kHz oppure int16."""
if audio16k.dtype != np.float32:
audio16k = audio16k.astype(np.float32) / 32768.0
if self._backend is None:
self.load()
if self._backend == "mlx":
import mlx_whisper
res = mlx_whisper.transcribe(audio16k, path_or_hf_repo=self.model, language="it", fp16=True)
return str(res.get("text", "")).strip()
segments, _ = self._fw.transcribe(audio16k, language="it", beam_size=1, vad_filter=True)
return " ".join(s.text for s in segments).strip()
+211
View File
@@ -0,0 +1,211 @@
"""Accesso Telegram con l'account dell'utente (Telethon): login e chiamate sincrone per i plugin."""
from __future__ import annotations
import asyncio
import json
import threading
from pathlib import Path
from typing import Any, Awaitable, Callable
from avatar.settings import DATA_DIR, Settings
SESSION = DATA_DIR / "telegram" # Telethon aggiunge .session
LOGIN_STATE = DATA_DIR / "telegram_login.json"
_lock = threading.Lock()
def credentials() -> tuple[int, str]:
s = Settings()
api_id = str(s.get("telegram_api_id") or "").strip()
api_hash = s.get_secret("telegram_api_hash")
if not api_id.isdigit() or not api_hash:
raise RuntimeError("Telegram non configurato: inserisci api id e api hash in Motore e Voce (da my.telegram.org).")
return int(api_id), api_hash
def run(fn: Callable[[Any], Awaitable[Any]], require_auth: bool = True) -> Any:
"""Esegue `fn(client)` in un event loop dedicato, con connessione aperta e chiusa ogni volta."""
from telethon import TelegramClient
api_id, api_hash = credentials()
async def main():
client = TelegramClient(str(SESSION), api_id, api_hash)
await client.connect()
try:
if require_auth and not await client.is_user_authorized():
raise RuntimeError("Telegram: accesso non ancora effettuato. Vai in Motore e Voce e completa l'accesso.")
return await fn(client)
finally:
await client.disconnect()
with _lock:
loop = asyncio.new_event_loop()
try:
return loop.run_until_complete(main())
finally:
loop.close()
# ── Login in due passi (dalla finestra impostazioni) ─────────────────────────
def send_code(phone: str) -> str:
phone = phone.strip().replace(" ", "")
async def go(client):
prev = None
try:
prev = json.loads(LOGIN_STATE.read_text())
except Exception:
pass
if prev and prev.get("phone") == phone and prev.get("hash"):
# Seconda richiesta: Telegram lo rimanda con un altro metodo (SMS o chiamata).
try:
from telethon.tl.functions.auth import ResendCodeRequest
sent = await client(ResendCodeRequest(phone_number=phone, phone_code_hash=prev["hash"]))
LOGIN_STATE.write_text(json.dumps({"phone": phone, "hash": sent.phone_code_hash}))
via = type(sent.type).__name__.replace("SentCodeType", "")
return f"Codice reinviato via {via or 'altro metodo'}: controlla SMS o chiamata, poi inseriscilo qui."
except Exception as err:
LOGIN_STATE.unlink(missing_ok=True)
print(f"[telegram] resend fallito: {err}")
sent = await client.send_code_request(phone)
LOGIN_STATE.write_text(json.dumps({"phone": phone, "hash": sent.phone_code_hash}))
via = type(sent.type).__name__.replace("SentCodeType", "")
dove = "nell'app Telegram (messaggio dal mittente «Telegram»)" if via == "App" else f"via {via}"
return f"Codice inviato {dove}. Inseriscilo qui; se non arriva, premi di nuovo «Invia codice» per riceverlo con un altro metodo."
return run(go, require_auth=False)
def sign_in(code: str, password: str = "") -> str:
try:
st = json.loads(LOGIN_STATE.read_text())
except Exception:
st = None
if st is None or (not code.strip() and password):
if not password:
return "Prima premi «Invia codice» (oppure usa l'accesso con QR); se Telegram chiede la password, scrivila e premi Accedi."
async def only_password(client):
from telethon.errors import SessionPasswordNeededError # noqa: F401
await client.sign_in(password=password)
me = await client.get_me()
LOGIN_STATE.unlink(missing_ok=True)
return f"Accesso effettuato come {me.first_name or ''} {me.last_name or ''} (@{me.username or '—'})."
try:
return run(only_password, require_auth=False)
except Exception as err:
return f"Password non accettata: {err}"
async def go(client):
from telethon.errors import SessionPasswordNeededError
try:
await client.sign_in(st["phone"], code.strip().replace(" ", ""), phone_code_hash=st["hash"])
except SessionPasswordNeededError:
if not password:
return "Serve anche la password di verifica in due passaggi: inseriscila e premi di nuovo Accedi."
await client.sign_in(password=password)
me = await client.get_me()
LOGIN_STATE.unlink(missing_ok=True)
return f"Accesso effettuato come {me.first_name or ''} {me.last_name or ''} (@{me.username or '—'})."
return run(go, require_auth=False)
def status() -> str:
try:
credentials()
except RuntimeError as err:
return str(err)
if not Path(str(SESSION) + ".session").exists():
return "Credenziali presenti; accesso non ancora effettuato."
async def go(client):
if not await client.is_user_authorized():
return "Accesso non ancora effettuato."
me = await client.get_me()
return f"Collegato come {me.first_name or ''} {me.last_name or ''} (@{me.username or '—'})."
try:
return run(go, require_auth=False)
except Exception as err:
return f"Errore: {err}"
def logout() -> str:
async def go(client):
await client.log_out()
return "Disconnesso da Telegram."
try:
return run(go, require_auth=False)
finally:
Path(str(SESSION) + ".session").unlink(missing_ok=True)
# ── Accesso con codice QR (Impostazioni › Dispositivi › Collega dispositivo) ──
_qr_state: dict = {}
def qr_login(on_update: Callable[[str, bytes | None], None]) -> None:
"""Genera un QR, lo passa a `on_update(messaggio, png)` e attende la scansione (fino a 3 minuti).
Bloccante: chiamare da un thread. A ogni scadenza del QR ne genera uno nuovo."""
import io
import qrcode
from telethon import TelegramClient
from telethon.errors import SessionPasswordNeededError
api_id, api_hash = credentials()
async def main():
client = TelegramClient(str(SESSION), api_id, api_hash)
await client.connect()
try:
if await client.is_user_authorized():
me = await client.get_me()
on_update(f"Già collegato come {me.first_name or ''} (@{me.username or '—'}).", None)
return
pwd = _qr_state.get("password", "")
if pwd:
# QR già scansionato in precedenza: manca solo la password.
try:
await client.sign_in(password=pwd)
me = await client.get_me()
on_update(f"Accesso effettuato come {me.first_name or ''} {me.last_name or ''} (@{me.username or '—'}).", None)
return
except Exception:
pass
deadline = asyncio.get_event_loop().time() + 180
while asyncio.get_event_loop().time() < deadline:
try:
qr = await client.qr_login()
except SessionPasswordNeededError:
on_update("QR accettato: ora serve la password di verifica in due passaggi. Scrivila nel campo password e premi «Accedi».", None)
return
buf = io.BytesIO()
qrcode.make(qr.url).save(buf, format="PNG")
on_update("Inquadra il QR dal telefono: Telegram › Impostazioni › Dispositivi › Collega dispositivo desktop.", buf.getvalue())
try:
await qr.wait(timeout=max(5, (qr.expires - qr.expires.__class__.now(qr.expires.tzinfo)).total_seconds()))
break
except asyncio.TimeoutError:
continue
except SessionPasswordNeededError:
pwd = _qr_state.get("password", "")
if not pwd:
on_update("Serve la password di verifica in due passaggi: scrivila nel campo password e ripeti «Accesso con QR».", None)
return
await client.sign_in(password=pwd)
break
if await client.is_user_authorized():
me = await client.get_me()
on_update(f"Accesso effettuato come {me.first_name or ''} {me.last_name or ''} (@{me.username or '—'}).", None)
else:
on_update("QR scaduto senza scansione. Premi di nuovo «Accesso con QR».", None)
finally:
await client.disconnect()
with _lock:
loop = asyncio.new_event_loop()
try:
loop.run_until_complete(main())
finally:
loop.close()
+138
View File
@@ -0,0 +1,138 @@
"""Sintesi vocale: Kokoro in locale (voci italiane) oppure la voce di sistema di macOS."""
from __future__ import annotations
import re
import subprocess
import tempfile
import threading
from pathlib import Path
import numpy as np
KOKORO_VOICES = {"if_sara": "Sara (femminile)", "im_nicola": "Nicola (maschile)"}
OUT_RATE = 24000
def clean_for_speech(text: str) -> str:
"""Toglie Markdown, link ed emoji prima di leggere."""
t = re.sub(r"```[\s\S]*?```", " ", text)
t = re.sub(r"`([^`]+)`", r"\1", t)
t = re.sub(r"!\[[^\]]*\]\([^)]*\)", "", t)
t = re.sub(r"\[([^\]]+)\]\([^)]*\)", r"\1", t)
t = re.sub(r"https?://\S+", "", t)
t = re.sub(r"^#{1,6}\s+", "", t, flags=re.M)
t = re.sub(r"^\s*[-*+]\s+", "", t, flags=re.M)
t = re.sub(r"^\s*\d+\.\s+", "", t, flags=re.M)
t = re.sub(r"\*\*([^*]+)\*\*", r"\1", t)
t = re.sub(r"\*([^*]+)\*", r"\1", t)
t = re.sub(r"_([^_]+)_", r"\1", t)
t = re.sub(r"^\|.*\|$", "", t, flags=re.M)
t = re.sub(r"[*_#>|]", "", t)
t = re.sub(r"[\U0001F000-\U0001FAFF☀-➿️‍]", "", t)
return re.sub(r"\s+", " ", t).strip()
class SentenceSplitter:
"""Riceve testo incrementale e restituisce le frasi complete."""
_RE = re.compile(r'[^.!?\n]+[.!?]+["»”)]?\s+|[^\n]+\n+')
def __init__(self) -> None:
self._buf = ""
def push(self, delta: str) -> list[str]:
self._buf += delta
if self._buf.count("```") % 2 == 1:
return []
out, last = [], 0
for m in self._RE.finditer(self._buf):
out.append(m.group(0))
last = m.end()
self._buf = self._buf[last:]
return [s for s in (clean_for_speech(x) for x in out) if s]
def flush(self) -> list[str]:
rest, self._buf = clean_for_speech(self._buf), ""
return [rest] if rest else []
class KokoroVoice:
"""Kokoro-82M via il pacchetto `kokoro`; fonetica italiana con espeak-ng."""
def __init__(self, voice: str = "if_sara", speed: float = 1.0, on_status=None) -> None:
self.voice = voice
self.speed = speed
self._on_status = on_status or (lambda m: None)
self._pipe = None
self._lock = threading.Lock()
def load(self) -> None:
with self._lock:
if self._pipe is not None:
return
self._on_status("Carico la voce Kokoro…")
try:
from kokoro import KPipeline
self._pipe = KPipeline(lang_code="i", repo_id="hexgrad/Kokoro-82M")
for _ in self._pipe("ciao", voice=self.voice, speed=self.speed):
pass
finally:
self._on_status(None)
def synthesize(self, text: str) -> np.ndarray:
if self._pipe is None:
self.load()
chunks = []
with self._lock:
for _, _, audio in self._pipe(text, voice=self.voice, speed=self.speed):
if audio is not None:
arr = audio.detach().cpu().numpy() if hasattr(audio, "detach") else np.asarray(audio)
chunks.append(arr.astype(np.float32).reshape(-1))
if not chunks:
return np.zeros(0, dtype=np.float32)
return np.concatenate(chunks)
class SystemVoice:
"""Voce di macOS tramite `say`, resa in un file audio e riprodotta dall'app."""
def __init__(self, voice: str = "") -> None:
self.voice = voice
@staticmethod
def list_voices() -> list[str]:
try:
out = subprocess.run(["say", "-v", "?"], capture_output=True, text=True, timeout=10).stdout
except Exception:
return []
names = []
for line in out.splitlines():
if "it_IT" in line:
names.append(line.split(" ")[0].strip())
return sorted(set(names))
def load(self) -> None:
pass
def synthesize(self, text: str) -> np.ndarray:
import soundfile as sf
voice = self.voice or next(iter(v for v in self.list_voices() if v.startswith("Alice")), "") or ""
with tempfile.TemporaryDirectory() as tmp:
path = Path(tmp) / "say.wav"
cmd = ["say", "-o", str(path), "--data-format=LEI16@24000"]
if voice:
cmd += ["-v", voice]
cmd.append(text)
subprocess.run(cmd, check=True, timeout=120)
data, rate = sf.read(str(path), dtype="float32")
if data.ndim > 1:
data = data[:, 0]
if rate != OUT_RATE:
data = np.interp(np.arange(0, data.size, rate / OUT_RATE), np.arange(data.size), data).astype(np.float32)
return data
def make_voice(settings, on_status=None):
if settings.get("tts_engine") == "system":
return SystemVoice(settings.get("system_voice", ""))
return KokoroVoice(settings.get("kokoro_voice", "if_sara"), on_status=on_status)
+25
View File
@@ -0,0 +1,25 @@
"""Ricerca web con Brave Search (per il motore locale)."""
from __future__ import annotations
import re
import requests
def brave_search(query: str, api_key: str, count: int = 6) -> list[dict]:
res = requests.get(
"https://api.search.brave.com/res/v1/web/search",
params={"q": query, "count": count, "country": "IT", "search_lang": "it"},
headers={"Accept": "application/json", "X-Subscription-Token": api_key},
timeout=15,
)
res.raise_for_status()
hits = []
for r in (res.json().get("web") or {}).get("results") or []:
if r.get("url"):
hits.append({
"title": r.get("title") or r["url"],
"url": r["url"],
"description": re.sub(r"<[^>]+>", "", r.get("description") or ""),
})
return hits
+184
View File
@@ -0,0 +1,184 @@
"""Gestione del ponte WhatsApp (processo Node con Baileys) e client della sua API locale."""
from __future__ import annotations
import os
import shutil
import subprocess
import threading
import time
from pathlib import Path
import requests
from avatar.settings import BASE_DIR, DATA_DIR
BRIDGE_DIR = BASE_DIR / "whatsapp_bridge"
WA_DATA = DATA_DIR / "whatsapp"
AUTH_DIR = WA_DATA / "auth"
PORT = 8790
_proc: subprocess.Popen | None = None
_lock = threading.Lock()
def node_path() -> str | None:
for cand in (shutil.which("node"), "/opt/homebrew/bin/node", "/usr/local/bin/node", os.path.expanduser("~/.nvm/current/bin/node")):
if cand and Path(cand).exists():
return cand
return None
def is_linked() -> bool:
return (AUTH_DIR / "creds.json").exists()
def running() -> bool:
return _proc is not None and _proc.poll() is None
def sync_contacts() -> int:
"""Copia nomi e numeri dalla rubrica del Mac nella tabella dei contatti WhatsApp e rinomina le chat già registrate."""
import re
import sqlite3
import Contacts
store = Contacts.CNContactStore.alloc().init()
keys = [Contacts.CNContactGivenNameKey, Contacts.CNContactFamilyNameKey, Contacts.CNContactOrganizationNameKey, Contacts.CNContactPhoneNumbersKey]
req = Contacts.CNContactFetchRequest.alloc().initWithKeysToFetch_(keys)
pairs: list[tuple[str, str]] = []
def visit(contact, stop):
name = f"{contact.givenName()} {contact.familyName()}".strip() or contact.organizationName()
if not name:
return
for ph in contact.phoneNumbers():
digits = re.sub(r"\D", "", ph.value().stringValue())
if digits.startswith("00"):
digits = digits[2:]
if len(digits) == 10 and digits.startswith("3"):
digits = "39" + digits
if len(digits) >= 10:
pairs.append((f"{digits}@s.whatsapp.net", name))
ok, err = store.enumerateContactsWithFetchRequest_error_usingBlock_(req, None, visit)
db = WA_DATA / "index.sqlite"
con = sqlite3.connect(db, timeout=10)
try:
con.execute("CREATE TABLE IF NOT EXISTS wa_contacts (jid TEXT PRIMARY KEY, name TEXT, notify TEXT)")
con.executemany("INSERT INTO wa_contacts (jid, name, notify) VALUES (?, ?, NULL) ON CONFLICT(jid) DO UPDATE SET name = excluded.name", pairs)
# Rinomina le chat registrate come numero
con.execute("UPDATE messages SET chat = (SELECT name FROM wa_contacts c WHERE c.jid = messages.jid) WHERE jid IN (SELECT jid FROM wa_contacts) AND chat GLOB '[0-9]*'")
con.execute("UPDATE messages SET sender = (SELECT name FROM wa_contacts c WHERE c.jid = messages.jid) WHERE from_me = 0 AND jid IN (SELECT jid FROM wa_contacts) AND sender GLOB '[0-9]*'")
con.commit()
finally:
con.close()
return len(pairs)
def start(me_name: str = "") -> str:
"""Avvia il ponte se non è già attivo. Restituisce un messaggio di stato."""
global _proc
with _lock:
if running():
return "Ponte WhatsApp già attivo."
node = node_path()
if not node:
return "Node.js non trovato: installa Node (brew install node) per usare WhatsApp in tempo reale."
if not (BRIDGE_DIR / "node_modules").exists():
return "Dipendenze del ponte mancanti: esegui `npm install` in whatsapp_bridge/."
WA_DATA.mkdir(parents=True, exist_ok=True)
log = open(WA_DATA / "bridge.log", "a")
_proc = subprocess.Popen([node, str(BRIDGE_DIR / "index.js"), "--data", str(WA_DATA), "--port", str(PORT), "--me", me_name or "io"],
stdout=log, stderr=subprocess.STDOUT, cwd=str(BRIDGE_DIR))
threading.Thread(target=lambda: _safe_sync(), daemon=True).start()
for _ in range(30):
time.sleep(0.5)
try:
status()
return "Ponte WhatsApp avviato."
except Exception:
if _proc.poll() is not None:
return "Il ponte WhatsApp si è chiuso subito: controlla data/whatsapp/bridge.log."
return "Il ponte WhatsApp non risponde."
def _safe_sync() -> None:
try:
time.sleep(3)
sync_contacts()
except Exception as err:
print(f"[whatsapp] sincronizzazione contatti fallita: {err}")
def stop() -> None:
global _proc
with _lock:
if _proc and _proc.poll() is None:
_proc.terminate()
try:
_proc.wait(5)
except Exception:
_proc.kill()
_proc = None
def _get(path: str, **params) -> dict:
r = requests.get(f"http://127.0.0.1:{PORT}{path}", params=params, timeout=8)
r.raise_for_status()
return r.json()
def status() -> dict:
return _get("/status")
def status_text() -> str:
if not running():
return "Ponte non avviato." + (" Sessione salvata: si collegherà all'avvio." if is_linked() else "")
try:
st = status()
except Exception as err:
return f"Ponte non raggiungibile: {err}"
c = st.get("connection")
if c == "open":
me = st.get("me") or {}
return f"WhatsApp collegato come {me.get('name') or me.get('id') or '?'} · messaggi ricevuti in questa sessione: {st.get('received', 0)}"
if c == "qr":
return "In attesa della scansione del QR."
return f"Non collegato. {st.get('error') or ''}".strip()
def qr() -> str | None:
try:
return status().get("qr")
except Exception:
return None
def qr_png() -> bytes | None:
code = qr()
if not code:
return None
import io
import qrcode
buf = io.BytesIO()
qrcode.make(code).save(buf, format="PNG")
return buf.getvalue()
def logout() -> str:
try:
_get("/logout")
except Exception:
pass
shutil.rmtree(AUTH_DIR, ignore_errors=True)
return "WhatsApp scollegato."
def resolve(to: str) -> dict:
return _get("/resolve", to=to)
def send(to: str, text: str) -> dict:
r = requests.post(f"http://127.0.0.1:{PORT}/send", json={"to": to, "text": text}, timeout=30)
if r.status_code != 200:
raise RuntimeError(r.json().get("error", r.text))
return r.json()
+215
View File
@@ -0,0 +1,215 @@
import * as THREE from 'three';
import { GLTFLoader } from './vendor/GLTFLoader.js';
const COLORS = { LISTENING: '#00ff88', THINKING: '#ffcc00', PROCESSING: '#ffcc00', SPEAKING: '#ff6b00', SLEEPING: '#3a8a9a', MUTED: '#ff3366' };
const canvas = document.getElementById('scene');
const renderer = new THREE.WebGLRenderer({ canvas, antialias: true, alpha: true });
renderer.setPixelRatio(Math.min(2, window.devicePixelRatio || 1));
renderer.outputColorSpace = THREE.SRGBColorSpace;
renderer.toneMapping = THREE.ACESFilmicToneMapping;
const scene = new THREE.Scene();
scene.background = new THREE.Color('#00060a');
const camera = new THREE.PerspectiveCamera(28, 1, 0.05, 50);
const hemi = new THREE.HemisphereLight('#bfe8ff', '#0a1a2a', 1.2);
const key = new THREE.DirectionalLight('#ffffff', 2.2); key.position.set(1.2, 2.4, 2.0);
const fill = new THREE.DirectionalLight('#9fd8ff', 0.8); fill.position.set(-1.5, 1.0, 1.5);
const rim = new THREE.DirectionalLight('#00d4ff', 1.6); rim.position.set(-2, 1.5, -2);
scene.add(hemi, key, fill, rim);
const state = { name: 'LISTENING', level: 0, open: 0, width: 0, framing: 'bust' };
const smooth = { open: 0, width: 0, nod: 0, yaw: 0, pitch: 0, glanceX: 0, glanceY: 0, glanceUntil: 0, light: 1 };
let model = null, mixer = null, head = null;
const eyes = [];
const restQ = new Map(); // orientamento nello spazio della posa neutra, per testa e occhi
const headPos = new THREE.Vector3(0, 1.6, 0);
const morphMeshes = [];
const mouth = { hasShapes: false, map: {} };
const blink = { next: 2, until: 0 };
const levels = new Array(36).fill(0);
const clock = new THREE.Clock();
let modelHeight = 1.7;
const modelCenter = new THREE.Vector3();
function resize() {
const w = window.innerWidth || canvas.clientWidth, h = window.innerHeight || canvas.clientHeight;
renderer.setSize(w, h, false);
camera.aspect = w / h;
camera.updateProjectionMatrix();
}
window.addEventListener('resize', resize);
resize();
function findMorphs(root) {
root.traverse((o) => {
if (o.isMesh && o.morphTargetDictionary && o.morphTargetInfluences) morphMeshes.push(o);
});
const names = new Set();
for (const m of morphMeshes) for (const k of Object.keys(m.morphTargetDictionary)) names.add(k);
const pick = (...cands) => cands.find((c) => names.has(c));
mouth.map = {
open: pick('jawOpen', 'JawOpen', 'mouthOpen', 'MouthOpen', 'viseme_aa', 'mouth_open', 'A'),
aa: pick('viseme_aa'),
smile: pick('mouthSmileLeft', 'mouthSmile', 'mouthSmile_L'),
smileR: pick('mouthSmileRight', 'mouthSmile_R'),
wide: pick('viseme_I', 'viseme_E'),
round: pick('viseme_O', 'mouthFunnel', 'mouthPucker', 'O'),
pucker: pick('mouthPucker'),
blinkL: pick('eyeBlinkLeft', 'eyesClosed'),
blinkR: pick('eyeBlinkRight'),
browUp: pick('browInnerUp'),
close: pick('mouthClose'),
};
mouth.hasShapes = Boolean(mouth.map.open);
document.getElementById('hint').textContent = (mouth.hasShapes ? 'labiale: attiva' : 'labiale: modello senza blend shape') + ' · clic: busto / figura intera';
}
function setMorph(name, value) {
if (!name) return;
for (const m of morphMeshes) {
const i = m.morphTargetDictionary[name];
if (i !== undefined) m.morphTargetInfluences[i] = value;
}
}
function measure() {
const box = new THREE.Box3().setFromObject(model);
modelHeight = Math.max(0.1, box.max.y - box.min.y);
box.getCenter(modelCenter);
}
function frame() {
if (!model) return;
if (head) head.getWorldPosition(headPos);
const h = modelHeight;
if (state.framing === 'bust') {
camera.position.set(headPos.x, headPos.y + h * 0.02, headPos.z + h * 0.56);
camera.lookAt(headPos.x, headPos.y - h * 0.015, headPos.z);
} else {
camera.position.set(modelCenter.x, modelCenter.y + h * 0.05, modelCenter.z + h * 1.9);
camera.lookAt(modelCenter.x, modelCenter.y, modelCenter.z);
}
}
canvas.addEventListener('click', () => { state.framing = state.framing === 'bust' ? 'full' : 'bust'; frame(); });
const MODEL = new URLSearchParams(location.search).get('model') || 'allegra_2';
new GLTFLoader().load(`./models/${encodeURIComponent(MODEL)}.glb`, (gltf) => {
model = gltf.scene;
scene.add(model);
model.traverse((o) => {
if (o.isBone) {
if (o.name === 'Head') head = o;
else if (o.name === 'LeftEye' || o.name === 'RightEye') eyes.push(o);
}
if (o.isMesh) o.frustumCulled = false;
});
findMorphs(model);
model.updateMatrixWorld(true);
for (const b of [head, ...eyes]) if (b) restQ.set(b, b.getWorldQuaternion(new THREE.Quaternion()));
let clipTracks = 0;
if (gltf.animations.length) {
mixer = new THREE.AnimationMixer(model);
const clip = gltf.animations[0].clone();
// Testa e occhi li orientiamo noi; il bacino resta fermo così il corpo guarda avanti.
clip.tracks = clip.tracks.filter((tr) => !/^(Head|Neck|Hips|LeftEye|RightEye)\./.test(tr.name));
clipTracks = clip.tracks.length;
const action = mixer.clipAction(clip);
action.setLoop(THREE.LoopRepeat); action.play();
}
document.getElementById('msg').remove();
measure();
if (head) head.getWorldPosition(headPos);
console.log('modello', MODEL, ': altezza', modelHeight.toFixed(3), 'testa', headPos.toArray().map((v) => v.toFixed(2)).join(','), 'blend shape', mouth.hasShapes, 'occhi', eyes.length, 'tracce', clipTracks);
frame();
}, undefined, (err) => {
document.getElementById('msg').textContent = 'Modello non caricato: ' + (err.message || err);
});
const qOffset = new THREE.Quaternion(), qParent = new THREE.Quaternion(), euler = new THREE.Euler();
/** Orienta un osso nello spazio: posa neutra + rotazione (pitch, yaw) in assi mondo. */
function aimBone(bone, pitch, yaw, roll) {
const rest = restQ.get(bone);
if (!rest || !bone.parent) return;
bone.parent.updateWorldMatrix(true, false);
bone.parent.getWorldQuaternion(qParent).invert();
euler.set(pitch, yaw, roll || 0, 'YXZ');
qOffset.setFromEuler(euler);
bone.quaternion.copy(qParent).multiply(qOffset).multiply(rest);
}
function render() {
requestAnimationFrame(render);
const dt = Math.min(0.05, clock.getDelta());
const t = clock.elapsedTime;
if (mixer) mixer.update(dt);
// Posa per stato (pitch positivo = guarda in basso)
let targetPitch = 0, targetYaw = 0, light = 1;
if (state.name === 'THINKING' || state.name === 'PROCESSING') { targetPitch = -0.14; targetYaw = 0.28 + Math.sin(t * 0.7) * 0.05; }
else if (state.name === 'SLEEPING') { targetPitch = 0.32; light = 0.45; }
else if (state.name === 'LISTENING') { targetYaw = Math.sin(t * 0.5) * 0.05; targetPitch = Math.sin(t * 0.8) * 0.02; }
else if (state.name === 'SPEAKING') { targetYaw = Math.sin(t * 0.9) * 0.04; }
if (t < smooth.glanceUntil) { targetYaw += smooth.glanceX; targetPitch += smooth.glanceY; }
const k = Math.min(1, dt * 4);
smooth.yaw += (targetYaw - smooth.yaw) * k;
smooth.pitch += (targetPitch - smooth.pitch) * k;
smooth.light += (light - smooth.light) * k;
hemi.intensity = 1.2 * smooth.light; key.intensity = 2.2 * smooth.light; fill.intensity = 0.8 * smooth.light;
const nodTarget = state.name === 'SPEAKING' ? state.level * 0.07 : 0;
smooth.nod += (nodTarget - smooth.nod) * Math.min(1, dt * 12);
if (head) aimBone(head, smooth.pitch + smooth.nod, smooth.yaw, Math.sin(t * 0.9) * 0.012);
for (const e of eyes) aimBone(e, smooth.pitch * 0.5 + Math.sin(t * 0.37) * 0.03, smooth.yaw * 0.6 + Math.sin(t * 0.23) * 0.04, 0);
// Bocca e viso
const openTarget = state.name === 'SPEAKING' ? state.open : 0;
smooth.open += (openTarget - smooth.open) * Math.min(1, dt * 18);
smooth.width += ((state.name === 'SPEAKING' ? state.width : 0) - smooth.width) * Math.min(1, dt * 12);
if (mouth.hasShapes) {
const o = Math.min(1, smooth.open), w = smooth.width;
setMorph(mouth.map.open, o * 0.45);
setMorph(mouth.map.aa, o * 0.55 * (1 - Math.abs(w) * 0.5));
setMorph(mouth.map.wide, Math.max(0, w) * o * 0.6);
setMorph(mouth.map.round, Math.max(0, -w) * o * 0.7);
setMorph(mouth.map.pucker, Math.max(0, -w) * o * 0.2);
const smileBase = state.name === 'LISTENING' ? 0.07 : 0.03;
setMorph(mouth.map.smile, smileBase + Math.max(0, w) * 0.2);
setMorph(mouth.map.smileR, smileBase + Math.max(0, w) * 0.2);
setMorph(mouth.map.close, 0); // usato da solo deforma le labbra: mai a riposo
setMorph(mouth.map.browUp, state.name === 'THINKING' || state.name === 'PROCESSING' ? 0.5 : (state.name === 'LISTENING' ? 0.15 : 0));
if (t > blink.next) { blink.until = t + 0.13; blink.next = t + 2 + Math.random() * 4; }
const closed = state.name === 'SLEEPING' ? 1 : (t < blink.until ? 1 : 0);
setMorph(mouth.map.blinkL, closed); setMorph(mouth.map.blinkR, closed);
}
if (state.framing === 'bust' && head) frame();
renderer.render(scene, camera);
drawOverlay();
}
const wave = document.getElementById('wave'), wctx = wave.getContext('2d'), stateEl = document.getElementById('state');
let dispLevel = 0;
function drawOverlay() {
dispLevel += (state.level - dispLevel) * 0.3;
levels.push(dispLevel); levels.shift();
const col = COLORS[state.name] || '#00d4ff';
stateEl.style.color = col; stateEl.style.textShadow = `0 0 8px ${col}88`;
stateEl.textContent = '● ' + state.name;
wctx.clearRect(0, 0, wave.width, wave.height);
const n = levels.length, bw = wave.width / n;
for (let i = 0; i < n; i++) {
const h = 3 + levels[i] * (wave.height - 4) + Math.sin(i * 0.7 + performance.now() / 600) * 1.5;
wctx.fillStyle = col; wctx.globalAlpha = 0.35 + 0.65 * (i / n);
wctx.fillRect(i * bw + 1, (wave.height - h) / 2, bw - 3, h);
}
wctx.globalAlpha = 1;
}
window.avatar3d = {
update(level, open, width, name) { state.level = level; state.open = open; state.width = width; if (name) state.name = name; },
setState(name) { state.name = name; },
glance(dx, dy, hold) { smooth.glanceX = dx * 0.3; smooth.glanceY = -dy * 0.2; smooth.glanceUntil = clock.elapsedTime + (hold || 1.1); },
setFraming(f) { state.framing = f; frame(); },
status() { return { loaded: Boolean(model), animation: Boolean(mixer), head: Boolean(head), blendShapes: mouth.hasShapes, morphs: morphMeshes.length, eyes: eyes.length }; },
};
render();
+27
View File
@@ -0,0 +1,27 @@
<!doctype html>
<html lang="it">
<head>
<meta charset="utf-8" />
<title>Avatar 3D</title>
<style>
html, body { margin: 0; height: 100%; background: #00060a; overflow: hidden; font-family: Menlo, monospace; }
#scene { position: absolute; left: 0; top: 0; width: 100%; height: 100%; display: block; }
#overlay { position: absolute; left: 0; right: 0; bottom: 14px; text-align: center; pointer-events: none; }
#state { font-size: 15px; font-weight: bold; letter-spacing: 2px; color: #00d4ff; text-shadow: 0 0 8px rgba(0,212,255,.5); }
#wave { display: block; margin: 8px auto 0; }
#hint { position: absolute; top: 10px; right: 12px; font-size: 11px; color: #3a8a9a; }
#msg { position: absolute; top: 45%; left: 0; right: 0; text-align: center; color: #79c0d2; font-size: 14px; }
</style>
<script type="importmap">{ "imports": { "three": "./vendor/three.module.js" } }</script>
</head>
<body>
<canvas id="scene"></canvas>
<div id="msg">Carico l'avatar…</div>
<div id="hint">clic: busto / figura intera</div>
<div id="overlay">
<div id="state">● LISTENING</div>
<canvas id="wave" width="360" height="40"></canvas>
</div>
<script type="module" src="./app.js"></script>
</body>
</html>
+2
View File
@@ -0,0 +1,2 @@
Metti qui i modelli GLB degli avatar (per esempio esportati da Avaturn con i blend shape).
Il nome del file diventa il nome nel menu "Avatar 3D" delle impostazioni.
File diff suppressed because it is too large. Load diff
+4727
View File
File diff suppressed because it is too large. Load diff
+54571
View File
File diff suppressed because it is too large. Load diff
BIN
View File
Binary file not shown.

After

Width:  |  Height:  |  Size: 71 KiB

View File
Whitespace-only changes.
+421
View File
@@ -0,0 +1,421 @@
"""
core/audio_devices.py — pick which microphone and which speakers JARVIS uses.
WHY
Both audio streams in main.py were opened without a `device=` argument, so
they always took whatever the operating system called "default". On a laptop
with a built-in mic, a webcam mic and a headset that is a coin toss — and on
Windows the default *moves on its own* the moment you plug a headset in.
"JARVIS can't hear me" almost always means "JARVIS is listening to the
monitor's microphone".
WHY NAMES, NOT INDICES
sounddevice identifies devices by integer index, and those indices shift
whenever a device appears or disappears. Storing index 3 means that after
unplugging a USB interface the saved setting silently points at something
else. We store the device *name* and resolve it to an index at open time.
WHY THIS IS CACHED
`sd.query_devices()` talks to the host audio API and can take a few hundred
milliseconds on a Windows machine with many endpoints. Mark LV learned this
lesson the expensive way — a 2.1-second `openwakeword` import on the Qt
thread made the settings drawer look like it was broken. So the list is
fetched once on a background thread at startup and served from cache.
"""
from __future__ import annotations
import threading
import time
# The label shown for "let the OS decide", and the value stored in config for
# it. Empty string, so an untouched install and a deliberately-default install
# are the same thing — nothing changes for anyone who never opens the picker.
DEFAULT_LABEL = "System default"
DEFAULT_VALUE = ""
_cache: dict[str, list[str]] | None = None
_cache_lock = threading.Lock()
# Which host API each direction settled on, so resolve() opens the same endpoint
# the picker listed. Filled in by _query().
_chosen_api: dict = {"input": None, "output": None}
# ── Why the raw list is unusable, and what is filtered out ───────────────────
#
# `sd.query_devices()` returns one entry per (device × host API), not one per
# device. Measured on a normal Windows machine: 41 entries for what the Windows
# sound settings show as 4 microphones and 4 speakers. The same Realtek
# microphone appears four times — once each under MME, DirectSound, WASAPI and
# WDM-KS — and none of the four is labelled to say which is which.
#
# Handing that to a person is not a choice, it is a quiz. So the list is reduced
# the way the operating system's own settings panel does it:
#
# 1. ONE host API per direction — chosen by measurement, not by reasoning.
# See the preference note below.
# 2. No pseudo-devices. "Microsoft Sound Mapper", "Primary Sound Driver",
# ALSA's "default"/"sysdefault"/"dmix" are aliases for "whatever the OS
# picks" — which is precisely the "System default" entry already at the top
# of the list. Offering them again as if they were hardware is noise.
# 3. No zero-channel or unnamed entries (WASAPI reports one of each).
# 4. Deduplicated by name.
#
# Nothing is hidden that a person could actually want: the same hardware is
# still there, listed once, under the name their operating system uses for it.
# ── Preference order, which is only a starting point ─────────────────────────
#
# Two earlier versions of this file were wrong in the same way: they decided
# which host API to use by reasoning about it instead of measuring it.
#
# v1 ranked APIs by how clean their device names were and picked WASAPI.
# WASAPI in shared mode does not resample — the hardware runs at 48 kHz,
# this app streams 16 kHz in and 24 kHz out, and every open failed with
# "Invalid sample rate". The picker looked right and did nothing.
#
# v2 added a rate check and picked DirectSound, which passes that check on
# both sides. PortAudio's DirectSound *output* is a silent sink: the
# stream opens, every write returns success in ~0 ms, and nothing is ever
# heard. Same failure, one layer deeper.
#
# So this list is a preference, not a promise. Which API actually gets used is
# decided below by _usable() (open it for real) and _transport_works() (does
# audio actually move), per direction. On this platform that lands on
# DirectSound for the microphone and MME for the speakers — a split that no
# amount of reasoning would have produced.
_PREFERRED_APIS = {
"Windows": ("directsound", "mme", "wasapi"),
# macOS has only Core Audio, so there is nothing to disambiguate.
"Darwin": ("core audio",),
# PulseAudio/PipeWire present one clean endpoint per device; raw ALSA
# presents dozens of routing permutations of the same card.
"Linux": ("pulse", "pipewire", "jack", "alsa"),
}
# ── "It opens" is not "it works" ─────────────────────────────────────────────
#
# Opening a stream successfully proves nothing. Measured, writing 2.0 s of audio:
#
# device=None (MME) 2.02 s consumed in real time
# DirectSound, any output device 0.00 s swallowed instantly
#
# No flag or capability field reports this. The only thing that separates a real
# sink from a fake one is whether it consumes audio at the rate audio is
# consumed at — so that is what gets measured, once per host API per direction,
# on the background thread at startup, using silence.
#
# Each direction is probed **the way main.py actually uses it**. That is not a
# detail: DirectSound input passes a callback stream and fails a blocking read,
# so an earlier version of this probe rejected a microphone that works perfectly
# in the app. Probe the mode you ship, not the mode that is easier to write.
_PROBE_SECONDS = {"output": 0.6, "input": 0.35}
# Cache: {(api_name_or_None, kind): bool}
_probe_results: dict = {}
def _transport_works(idx: int, kind: str, api_key) -> bool:
"""Does this host API actually move audio, or only pretend to?
Probed once per API per direction and cached. Output writes silence, so the
probe is inaudible; input reads and discards."""
if api_key in _probe_results:
return _probe_results[api_key]
ok = False
try:
import sounddevice as sd
rate = _RATES.get(kind, 16000)
secs = _PROBE_SECONDS.get(kind, 0.5)
# Each direction is probed the way main.py actually uses it. That is not
# a detail: DirectSound input passes a callback stream and fails a
# blocking read, so probing the wrong mode rejected a microphone that
# works perfectly in the app.
if kind == "output":
# main.py writes with stream.write() — a real sink is rate-limited
# by the hardware clock, a fake one swallows the buffer instantly.
st = sd.RawOutputStream(samplerate=rate, channels=1, dtype="int16",
blocksize=1024, device=idx)
st.start()
t0 = time.monotonic()
st.write(bytes(int(rate * secs) * 2)) # silence — inaudible
elapsed = time.monotonic() - t0
st.stop(); st.close()
ok = elapsed > secs * 0.5
if not ok:
print(f"[Audio] output: host API reports success but moves no "
f"audio ({elapsed*1000:.0f} ms for {secs*1000:.0f} ms) "
f"— skipping it")
else:
# main.py reads through a callback — count what arrives.
frames = [0]
def _cb(indata, n, *_a):
frames[0] += n
st = sd.InputStream(samplerate=rate, channels=1, dtype="int16",
blocksize=1024, device=idx, callback=_cb)
st.start()
time.sleep(secs)
st.stop(); st.close()
ok = frames[0] > rate * secs * 0.3
if not ok:
print(f"[Audio] input: host API delivered {frames[0]} frames in "
f"{secs*1000:.0f} ms — skipping it")
except Exception as e:
print(f"[Audio] {kind} transport probe failed: {e}")
ok = False
_probe_results[api_key] = ok
return ok
def _display_name(name: str, devices) -> str:
"""MME truncates device names to 31 characters, so the API that actually
carries the audio may not be the one that can spell. If another host API
knows a longer name that starts with this one, show that instead — the user
reads 'Realtek HD Audio 2nd output (Realtek(R) Audio)' while the stream runs
on the endpoint called 'Realtek HD Audio 2nd output (Re'."""
if len(name) < 30:
return name
best = name
for dev in devices:
other = (dev.get("name") or "").strip()
if len(other) > len(best) and other.startswith(name):
best = other
return best
# The rates the app opens its streams at. Defaults match main.py; main.py calls
# configure() at startup with its own constants so the two can never drift apart
# and silently reintroduce the bug above.
_RATES = {"input": 16000, "output": 24000}
def configure(input_rate: int, output_rate: int) -> None:
"""Tell this module the sample rates the audio streams will use, so the
picker can rule out devices that cannot be opened at them.
Drops any cached list: which devices are usable depends on the rate, so a
list built under the old rates would be stale."""
global _cache
_RATES["input"] = int(input_rate)
_RATES["output"] = int(output_rate)
with _cache_lock:
_cache = None
def _usable(idx: int, kind: str) -> bool:
"""Can this device actually be opened at the rate we need?
Deliberately opens a real stream rather than asking
`check_output_settings`, because that function lies: it passed for an MME
endpoint that then failed to open with "The specified format is not
supported or cannot be translated" [MME error 32]. Opening and immediately
closing costs milliseconds and is the only answer that holds."""
st = None
try:
import sounddevice as sd
rate = _RATES.get(kind, 16000)
if kind == "input":
st = sd.InputStream(samplerate=rate, channels=1, dtype="int16",
blocksize=1024, device=idx,
callback=lambda *_a: None)
else:
st = sd.RawOutputStream(samplerate=rate, channels=1, dtype="int16",
blocksize=1024, device=idx)
st.start()
return True
except Exception:
return False
finally:
if st is not None:
try:
st.stop(); st.close()
except Exception:
pass
# Aliases for "the default device" and internal routing endpoints. Matched
# case-insensitively as substrings against the device name.
_PSEUDO_DEVICES = (
"sound mapper", # Windows MME
"primary sound", # Windows DirectSound ("Primary Sound Capture Driver")
"sysdefault", # ALSA
"default", # ALSA / PulseAudio alias
"dmix", "dsnoop", # ALSA software mixing plugins
"surround", # ALSA channel-layout permutations of one card
"samplerate", "speexrate", "upmix", "vdownmix", "null",
)
def _is_pseudo(name: str) -> bool:
low = name.lower()
return any(tok in low for tok in _PSEUDO_DEVICES)
def _query() -> dict[str, list[str]]:
"""Return {'input': [names...], 'output': [names...]}. Never raises.
Only real, selectable devices — see the note above."""
out: dict[str, list[str]] = {"input": [], "output": []}
try:
import platform
import sounddevice as sd
devices = list(sd.query_devices())
try:
apis = [a.get("name", "") for a in sd.query_hostapis()]
except Exception:
apis = []
preferred = _PREFERRED_APIS.get(platform.system(), ())
def _collect(api_filter, kind) -> list[tuple[int, str]]:
"""(index, name) for named, non-pseudo devices on one side that can
be opened at the rate that side runs at."""
chan = "max_input_channels" if kind == "input" else "max_output_channels"
found, seen = [], set()
for idx, dev in enumerate(devices):
name = (dev.get("name") or "").strip()
if not name or _is_pseudo(name) or name in seen:
continue
if dev.get(chan, 0) <= 0:
continue
if api_filter is not None:
api = apis[dev["hostapi"]].lower() if dev.get("hostapi", -1) < len(apis) else ""
if api_filter not in api:
continue
if not _usable(idx, kind):
continue
seen.add(name)
found.append((idx, name))
return found
# Each direction picks its own host API. They are genuinely different
# problems — on Windows the microphone works on DirectSound while the
# speakers only work on MME — and a single global choice cannot be right
# for both.
for kind in ("input", "output"):
for api_filter in list(preferred) + [None]:
found = _collect(api_filter, kind)
if not found:
continue
# One probe per API per direction, cached, on this thread.
if not _transport_works(found[0][0], kind, (api_filter, kind)):
continue
_chosen_api[kind] = api_filter
out[kind] = [_display_name(n, devices) for _i, n in found]
break
if out[kind]:
print(f"[Audio] {kind}: using "
f"{_chosen_api[kind] or 'any host API'} "
f"({len(out[kind])} devices)")
return out
except Exception as e:
print(f"[Audio] Device enumeration failed: {e}")
return out
def prefetch() -> None:
"""Warm the cache on a background thread. Called once at startup so the
settings drawer never pays for enumeration on the Qt thread."""
def _work():
global _cache
result = _query()
with _cache_lock:
_cache = result
print(f"[Audio] {len(result['input'])} input / "
f"{len(result['output'])} output devices found")
threading.Thread(target=_work, daemon=True, name="audio-devices").start()
def list_devices(kind: str, refresh: bool = False) -> list[str]:
"""Device names for 'input' or 'output'. Falls back to a synchronous query
if the prefetch has not landed yet — correctness over the cache."""
global _cache
with _cache_lock:
cached = None if refresh else _cache
if cached is None:
cached = _query()
with _cache_lock:
_cache = cached
return list(cached.get(kind, []))
def resolve(name: str, kind: str):
"""Turn a saved device name into something sounddevice accepts.
Returns None for "system default" — which is also what we return when the
saved device is gone, because a missing headset must degrade to the built-in
speakers, not to a crash on startup.
Candidates are walked in the same host-API order the picker used, so a name
the user chose from the WASAPI list resolves to the WASAPI endpoint. Without
that ordering a full name would fall through to MME's truncated copy of the
same device — which happens to work, but means the setting quietly refers to
a different endpoint than the one on screen."""
wanted = (name or "").strip()
if not wanted or wanted == DEFAULT_LABEL:
return None
try:
import platform
import sounddevice as sd
devices = list(sd.query_devices())
try:
apis = [a.get("name", "") for a in sd.query_hostapis()]
except Exception:
apis = []
want_in = (kind == "input")
chan_key = "max_input_channels" if want_in else "max_output_channels"
def _candidates(api_filter):
for idx, dev in enumerate(devices):
if dev.get(chan_key, 0) <= 0:
continue
if api_filter is not None:
api = apis[dev["hostapi"]].lower() if dev.get("hostapi", -1) < len(apis) else ""
if api_filter not in api:
continue
yield idx, (dev.get("name") or "").strip()
# The API the picker settled on for this direction comes first — the
# endpoint that was listed must be the endpoint that gets opened, or the
# setting means something different from what it says. list_devices()
# populates it; calling it here is a no-op once the cache is warm.
list_devices(kind)
chosen = _chosen_api.get(kind)
orders = ([chosen] if chosen is not None else []) \
+ [a for a in _PREFERRED_APIS.get(platform.system(), ()) if a != chosen] \
+ [None]
# A candidate only counts if it can be opened at the rate this side runs
# at. The prefix match matters because the API that carries the audio is
# not always the one that can spell: MME truncates names to 31 characters
# while DirectSound and WASAPI do not, so the name shown in the picker
# can be longer than the name of the endpoint it actually opens.
for api_filter in orders:
partial = None
for idx, dev_name in _candidates(api_filter):
if dev_name == wanted:
if _usable(idx, kind):
return idx
continue
if partial is None and (dev_name.startswith(wanted[:24])
or wanted.startswith(dev_name[:24])):
if _usable(idx, kind):
partial = idx
if partial is not None:
return partial
print(f"[Audio] Saved {kind} device '{wanted}' cannot be opened at "
f"{_RATES.get(kind)} Hz on any host API — using system default")
return None
except Exception as e:
print(f"[Audio] resolve({kind}) failed: {e} — using system default")
return None
+715
View File
@@ -0,0 +1,715 @@
"""
Holographic AI head for the HUD centre — the thing that used to be a ring stack
with the assistant's name in the middle.
Design notes
------------
* **The face is real human geometry.** `core.avatar_mesh` builds the head around
MediaPipe's canonical face model, so eyelids, nostrils, lips and cheekbones
are measured anatomy rather than fitted curves. This renderer's whole job is
to light it, pose it and animate it.
* **Software rendered, on purpose.** Everything is QPainter, so there is no
OpenGL context, no shader compile, no GPU driver to disagree with us and no
new pip dependency. It looks the same on a gaming rig, a 2013 laptop, a VM
and a remote desktop session.
* **Lip-sync comes from the audio pipeline, not from the avatar.** `main.py`
already computes a real RMS level off the PCM (`_pcm_level`) for both the mic
and JARVIS's own output. The avatar just consumes that number, so there is no
second audio path to fall out of sync. The mouth only tracks the level while
JARVIS is *speaking* — during listening the same level drives the aura, so the
head never lip-syncs to the user's voice.
The renderer is theme-agnostic: `paint()` takes its colours as arguments, which
is what lets the HueWheel accent picker retint the avatar for free.
"""
from __future__ import annotations
import math
import random
import numpy as np
from PyQt6.QtCore import QLineF, QPointF, QRectF, Qt
from PyQt6.QtGui import QBrush, QColor, QPainter, QPen, QPolygonF, QRadialGradient
from core.avatar_mesh import JAW_MAX, JAW_PIVOT, get_head_mesh
# Perspective camera distance in head-half-heights. Large enough that the nose
# does not balloon, small enough to keep a sense of depth.
_CAM_D = 4.6
# Wireframe opacity buckets, so the whole lattice draws in a handful of batched
# drawLines() calls instead of one call per line.
_BUCKETS = 4
_MIN_ALPHA = 0.05
# Resolution of the surface-shading colour ramp. Banding across a filled facet
# is far more visible than banding in line alpha, so this is fine-grained — and
# being a lookup it costs nothing per face.
_LUT_N = 192
# How far the brows travel at full lift, in head-half-heights. Derived, not
# tuned: the brow-to-eye gap is 0.198 and a real raise covers about a third of
# it, then the drawn landmarks only carry half the rig weight.
_BROW_LIFT = 0.14
# Mouth timing, as time constants in seconds rather than per-frame fractions.
# A fixed per-frame lerp silently changes speed with the frame rate: the HUD
# runs at 60 Hz here, throttles its paint to 30, and drops to 20 when idle, so
# the same constant meant three different mouths. These do not.
#
# Shutting is the fastest of the three, and it is measured rather than chosen.
# A short closure occupies a single 20 ms schedule frame, so the mouth has one
# step to reach it: at 20 ms the jaw got a third of the way and the closure
# vanished, at 12 ms it arrives, and going below that changes nothing because
# the analysis window is then the limit, not the smoothing. Halving it doubled
# the closures the mouth visibly makes across a test paragraph, 5 of 21 to 10,
# with no loss of opening on the vowels. Only the return to rest, once talking
# has actually stopped, is leisurely.
_TAU_OPEN = 0.022 # jaw dropping toward a vowel
_TAU_SHUT = 0.012 # lips closing on a consonant, mid-word
_TAU_REST = 0.055 # settling back to rest after speech ends
_TAU_SHAPE = 0.018 # viseme openness following the schedule
# Only the microphone path needs a level floor: it has one coarse RMS and no way
# to tell speech from room tone. JARVIS's own voice arrives as a per-20 ms
# schedule whose silences are already silent, so it needs no floor and must not
# have one — a floor there swallows the gaps between words.
_MIC_FLOOR = 0.14
# How far below this voice's own loud level counts as a closure: -20 dB, which
# is what a stop consonant actually drops to. Expressed as a ratio so it holds
# at any speaker volume.
_CLOSE_FRAC = 0.10
def _rate(dt: float, tau: float) -> float:
"""Per-frame lerp factor for an exponential approach with time constant
`tau`. Frame-rate independent: the motion takes the same wall-clock time at
20, 30 or 60 fps, and a long frame catches up instead of stalling."""
return 1.0 - math.exp(-dt / tau)
def _c(col: QColor, a: float) -> QColor:
"""Copy of `col` at alpha `a` (0-255, clamped)."""
q = QColor(col)
q.setAlpha(int(max(0.0, min(255.0, a))))
return q
def _blend(bg: QColor, col: QColor, a: float) -> QColor:
"""`col` at alpha `a` pre-mixed onto `bg`, returned fully **opaque**.
Qt's raster engine has a fast path for opaque antialiased lines and a much
slower blended path for everything else — measured at 1.0 ms versus 2.9 ms
for the same 780 lines. The HUD paints a flat background behind the avatar,
so mixing the alpha in by hand is visually equivalent and three times cheaper.
"""
f = max(0.0, min(1.0, a / 255.0))
return QColor(int(bg.red() + (col.red() - bg.red()) * f),
int(bg.green() + (col.green() - bg.green()) * f),
int(bg.blue() + (col.blue() - bg.blue()) * f))
class HoloAvatar:
"""Animated holographic head. One instance per HUD canvas.
Lifecycle:
av = HoloAvatar()
av.step(dt, amp, speaking=..., muted=...) # once per tick
av.paint(painter, cx, cy, r, primary, accent, bg) # once per frame
"""
# Look: True paints a lit, solid head with a wireframe over it; False is a
# see-through glass wireframe. Flip here, or per instance.
shaded = True
def __init__(self) -> None:
mesh = get_head_mesh()
self._v0 = mesh["verts"]
self._n0 = mesh["normals"]
self._jaw = mesh["jaw"]
self._brow_w = mesh["brow"]
self._lips_w = mesh["lips"]
self._lip_c = mesh["lip_centre"]
self._fade = mesh["fade"]
self._f = mesh["faces"]
self._fgroup = mesh["face_group"]
self._fa, self._fb, self._fc = (self._f[:, i] for i in range(3))
self._e0 = mesh["edges"][:, 0]
self._e1 = mesh["edges"][:, 1]
self._lm = mesh["landmarks"]
# The inner-lip ring runs lower-lip left→right, then upper-lip back.
# Splitting it lets the upper arc anchor a strip of teeth, which is what
# keeps an open mouth from reading as a hole punched in the face.
lips_in = mesh["landmarks"]["lips_in"]
self._lip_up = np.concatenate([lips_in[10:], lips_in[:1]])
# Crown (+1.0) down to the bottom of the neck, in head-half-heights.
# Callers size the head to the room they have with this.
self.SPAN = mesh["span"][0] - mesh["span"][1]
self._lut_cache: list = []
self._lut_key = None
n = self._v0.shape[0]
self._v = np.empty((n, 3), dtype=np.float32)
self._t = 0.0
self._sway = 0.0 # integrated sway phase — see step()
self._yaw = 0.0
self._pitch = 0.0
self._mouth = 0.0 # 0..1 smoothed jaw opening
self._glow = 0.0 # 0..1 smoothed overall energy
self._scan = -1.6 # vertical position of the energy sweep
self._blink = 0.0 # 0 = open, 1 = shut
self._blink_at = 3.0
# ── expression ──────────────────────────────────────────────────────
# Speech is not just a moving jaw. Brows ride the loudness envelope,
# eyes widen with the brows, and the gaze flicks between fixation
# points — those three are what make it read as talking rather than as
# a puppet chewing.
self._amp_slow = 0.0
self._expr = 0.0
self._expr_tgt = 0.0
self._expr_at = 0.0
self._brow = 0.0 # smoothed brow lift, -0.4 .. 1.2
self._emph = 0.0 # syllable emphasis, drives the head nod
self._gaze = [0.0, 0.0]
self._gaze_tgt = [0.0, 0.0]
self._gaze_at = 0.0
# ── state expression ────────────────────────────────────────────────
# The face is the fastest status indicator in the app: you read a gaze
# before you read a word. Saccades orbit a bias that the assistant's
# state moves — eyes off to the side while it thinks, back on you while
# it listens, lids low while it sleeps.
self._gaze_bias = [0.0, 0.0]
self._bias_tgt = [0.0, 0.0]
self._bias_at = 0.0
self._lids = 1.0 # 1 = wide, 0 = shut; low while asleep
self._brow_bias = 0.0 # concentration pulls the brows down
self._glance = None # (dx, dy, until_t) — a deliberate look
# ── viseme ──────────────────────────────────────────────────────────
# Loudness alone only answers "how far open", which is why an RMS-driven
# mouth flaps rather than speaks. These two carry the *shape*: how open
# the jaw is for this sound, and whether the lips are spread (/i/) or
# rounded (/u/). They come from a formant read of the audio actually
# being played — see `_pcm_visemes` in main.py.
self._v_open = 1.0
self._v_wide = 0.0
self._wide = 0.0 # smoothed lip spread, -1 round .. +1 spread
self._v_peak = 0.18 # running estimate of this voice's loud level
# ── animation ───────────────────────────────────────────────────────────
def _mouth_step(self, dt: float, amp: float, live: bool,
v_open: float | None, v_level: float | None) -> None:
"""One increment of the jaw. Called once per viseme frame while JARVIS
speaks, once per rendered frame otherwise."""
if v_open is None:
shape = 1.0
else:
self._v_open += (v_open - self._v_open) * _rate(dt, _TAU_SHAPE)
shape = self._v_open
if v_level is None:
gated = max(0.0, (amp - _MIC_FLOOR) / (1.0 - _MIC_FLOOR))
drive = (gated ** 0.6) * (shape ** 0.75)
else:
# Speech RMS spends most of its time well below full scale, so the
# raw value alone would only ever half-open the jaw. Normalise it
# against a running estimate of this voice's own loud level rather
# than a constant: it then reads the same whether the user has the
# volume low or the model happens to be speaking softly.
self._v_peak = max(v_level, self._v_peak - dt * 0.55)
ref = max(0.18, self._v_peak)
# The floor is a fraction of this voice's own loud level, not a
# fixed number, so it means the same thing at any volume and in any
# language. A stop consonant drops 20 dB or more below the vowels
# around it, which is this ratio — so a real closure lands at
# exactly zero rather than at some small positive value the curve
# would otherwise lift back up. That lift is what kept the mouth
# from ever quite shutting between words.
q = (v_level - _CLOSE_FRAC * ref) / (ref * (1.0 - _CLOSE_FRAC))
drive = max(0.0, min(1.0, q)) ** 0.85 * (shape ** 0.75)
target = min(1.0, drive) if live else 0.0
if target > self._mouth:
tau = _TAU_OPEN
elif live:
tau = _TAU_SHUT # mid-word: a consonant, and it must shut now
else:
tau = _TAU_REST # speech is over; settle, don't snap
self._mouth += (target - self._mouth) * _rate(dt, tau)
if self._mouth < 0.002:
self._mouth = 0.0
def step(self, dt: float, amp: float, speaking: bool = False,
muted: bool = False, state: str = "",
v_open: float | None = None, v_wide: float = 0.0,
v_level: float | None = None,
v_seq: list | None = None, v_hop: float = 0.02) -> None:
"""Advance the animation.
`amp` is the 0..1 display audio level. `v_open` / `v_wide` / `v_level`
are the viseme schedule's shape and true level for this instant; passing
None falls back to loudness-only articulation, which is what the
microphone path uses. `v_seq` is every schedule frame the last rendered
frame spanned, so no closure is lost when the paint rate drops.
"""
dt = max(0.001, min(0.10, float(dt)))
self._t += dt
t = self._t
amp = max(0.0, min(1.0, float(amp)))
live = speaking and not muted
# Idle sway. The phase is *integrated* rather than taken as
# sin(t * rate * speed): multiplying absolute time by a speed that
# changes when JARVIS starts or stops talking jumps the phase by
# t * rate * delta, which after a minute of uptime is several radians
# and visibly teleports the head the instant a sentence ends.
speed = (1.0 if not muted else 0.55) * (1.25 if live else 1.0)
self._sway += dt * speed
s = self._sway
self._yaw = 0.26 * math.sin(s * 0.31) + 0.09 * math.sin(s * 0.73 + 1.3)
self._pitch = (0.060 * math.sin(s * 0.23 + 0.7)
+ 0.024 * math.sin(s * 0.61))
# Mouth. Which level is driving it matters more than any rate here.
#
# `v_level` is this 20 ms frame's own RMS, taken from the very audio
# about to be heard, so its silences are real silences. `amp` is the
# waveform display's level, and that one is a *peak hold*: it keeps the
# loudest value it has seen and decays gently, on purpose, so the bars
# do not stutter between audio chunks. Driving a mouth from a peak hold
# is why the gaps between words never closed — the hold spans exactly
# the consonant it was supposed to reveal. So the schedule drives the
# jaw whenever there is one, and `amp` is left to the microphone path,
# which has nothing better.
# Advance the mouth once per *schedule* frame rather than once per
# rendered frame. A bilabial closure lasts around 40 ms — two frames of
# a 50 Hz schedule — and the HUD throttles its paint to 30 fps and to 20
# when idle. Point-sampling at 20 fps steps 50 ms at a time, so a whole
# closure can fall between two samples and simply never be seen; that is
# information loss no smoothing constant can recover. Sub-stepping costs
# a few float operations per frame and makes the mouth identical at 20,
# 30 and 60 fps.
# An empty list is meaningful and is not the same as None: it says a
# schedule is playing but this tick landed inside a frame already
# spoken. The mouth's clock is the schedule's, so the right thing then
# is to do nothing. Re-stepping the same frame — which is what a
# truthiness test here would do — advances the jaw twice for one 20 ms
# of audio, and at 60 fps that alone made the mouth behave differently
# than at 20.
if v_seq is not None:
for lv, op, _wd in v_seq:
self._mouth_step(v_hop, amp, live, op, lv)
else:
self._mouth_step(dt, amp, live, v_open, v_level)
# Syllable emphasis. Applied unconditionally: `_emph` decays to zero on
# its own once the mouth closes, whereas gating it on `live` deleted the
# whole offset in a single frame and snapped the head at sentence end.
self._emph += (self._mouth - self._emph) * _rate(
dt, 0.055 if self._mouth > self._emph else 0.32)
self._pitch -= self._emph * 0.028
self._yaw += 0.018 * math.sin(t * 1.7) * self._emph
# Loudness envelope, deliberately lazier than the mouth: brows track the
# shape of a phrase, not individual syllables.
env = amp if live else 0.0
self._amp_slow += (env - self._amp_slow) * _rate(
dt, 0.16 if env > self._amp_slow else 0.36)
if live:
if t >= self._expr_at:
self._expr_tgt = random.uniform(-0.35, 1.0)
self._expr_at = t + 1.1 + 2.0 * random.random()
else:
self._expr_tgt = 0.0
self._expr_at = t + 0.8
self._expr += (self._expr_tgt - self._expr) * 0.075
brow_t = 0.55 * self._amp_slow + 0.60 * self._expr + self._brow_bias
self._brow += (max(-0.4, min(1.2, brow_t)) - self._brow) * 0.20
# ── what the state does to the face ─────────────────────────────────
st = (state or "").upper()
thinking = st in ("THINKING", "PROCESSING")
asleep = st in ("SLEEPING", "STANDBY", "OFFLINE")
if thinking:
# People look away to think, and hold it. The direction re-rolls
# slowly so it reads as thought rather than as scanning.
if t >= self._bias_at:
self._bias_tgt = [random.choice((-1.0, 1.0)) * random.uniform(0.45, 0.8),
random.uniform(0.25, 0.55)]
self._bias_at = t + 1.4 + 1.6 * random.random()
brow_bias, lid_tgt = -0.28, 0.94
elif asleep:
self._bias_tgt = [0.0, -0.25]
brow_bias, lid_tgt = -0.05, 0.22
else:
# LISTENING / idle / speaking: eyes come back to the user.
self._bias_tgt = [0.0, 0.0]
self._bias_at = 0.0
brow_bias = 0.10 if st == "LISTENING" else 0.0
lid_tgt = 1.0
for i in (0, 1):
self._gaze_bias[i] += (self._bias_tgt[i] - self._gaze_bias[i]) * 0.06
self._lids += (lid_tgt - self._lids) * 0.08
self._brow_bias += (brow_bias - self._brow_bias) * 0.06
# Gaze: saccades are near-instant jumps between fixations, and they get
# more frequent when there is something to say. While thinking they slow
# right down — a darting eye reads as nervous, not thoughtful.
if t >= self._gaze_at:
reach = 0.9 if live else (0.35 if thinking else 0.55)
self._gaze_tgt = [random.uniform(-1.0, 1.0) * reach,
random.uniform(-1.0, 1.0) * reach * 0.55]
if live:
self._gaze_at = t + 0.55 + 1.7 * random.random()
elif thinking:
self._gaze_at = t + 1.8 + 2.4 * random.random()
else:
self._gaze_at = t + 1.3 + 2.8 * random.random()
# A deliberate glance (something appeared on screen) overrides the
# wandering for a moment, then hands control back.
if self._glance is not None:
gx, gy, until = self._glance
if t < until:
self._gaze_tgt = [gx, gy]
else:
self._glance = None
for i, b in enumerate(self._gaze_bias):
tgt = max(-1.0, min(1.0, self._gaze_tgt[i] + b))
self._gaze[i] += (tgt - self._gaze[i]) * 0.30
# Lips lead the jaw slightly in real speech, so they track a touch
# faster; they also relax to neutral the moment the voice stops.
wide_t = v_wide if (live and v_open is not None) else 0.0
self._wide += (max(-1.0, min(1.0, wide_t)) - self._wide) * _rate(dt, 0.030)
self._glow += ((0.0 if muted else amp) - self._glow) * (
0.35 if (0.0 if muted else amp) > self._glow else 0.10)
self._scan += dt * (0.55 + 1.5 * self._glow)
if self._scan > 1.35:
self._scan = -1.75
if self._blink > 0.0:
self._blink = max(0.0, self._blink - dt * 8.5)
elif t >= self._blink_at:
# Concentration suppresses blinking; a sleeping face has no need of
# it at all, since the lids are already down.
if asleep:
self._blink_at = t + 6.0
else:
self._blink = 1.0
gap = 5.5 if thinking else 3.4
self._blink_at = t + gap + 3.1 * random.random()
def glance(self, dx: float, dy: float, hold: float = 1.1) -> None:
"""Look deliberately somewhere for `hold` seconds, then wander again.
Used when something appears on screen: a face that looks at what just
showed up tells the user it landed, without a word being spoken.
"""
self._glance = (max(-1.0, min(1.0, float(dx))),
max(-1.0, min(1.0, float(dy))),
self._t + max(0.1, float(hold)))
# ── posing ──────────────────────────────────────────────────────────────
def _pose(self):
"""Jaw drop, brow lift and head rotation, applied to the real geometry."""
v = self._v
np.copyto(v, self._v0)
if self._brow > 0.004 or self._brow < -0.004:
# The brow-to-eye gap is 0.198 head-half-heights and a real raise
# moves a third of it. The old 0.045 — halved again by the landmark
# weights, which average 0.5 — worked out to six pixels on a 250 px
# head, which is to say invisible.
v[:, 1] += self._brow_w * (self._brow * _BROW_LIFT)
if abs(self._wide) > 0.01 and self._mouth > 0.0:
# Spread pulls the corners out and flattens the lips back; rounding
# draws them in and pushes them forward into a purse.
k = self._lips_w * (self._wide * self._mouth)
v[:, 0] += k * (v[:, 0] - self._lip_c[0]) * 0.55
v[:, 1] += k * (v[:, 1] - self._lip_c[1]) * 0.30
v[:, 2] -= k * 0.055
if self._mouth > 0.004:
px, py, pz = JAW_PIVOT
ang = self._jaw * (self._mouth * JAW_MAX)
ca, sa = np.cos(ang), np.sin(ang)
dy = v[:, 1] - py
dz = v[:, 2] - pz
v[:, 1] = py + dy * ca - dz * sa
v[:, 2] = pz + dy * sa + dz * ca
cy, sy = math.cos(self._yaw), math.sin(self._yaw)
cp, sp = math.cos(self._pitch), math.sin(self._pitch)
m = np.array([
[cy, 0.0, sy],
[sp * sy, cp, -sp * cy],
[-cp * sy, sp, cp * cy],
], dtype=np.float32)
return v @ m.T, self._n0 @ m.T
# ── rendering ───────────────────────────────────────────────────────────
def _lut(self, bg: QColor, primary: QColor):
"""Cached ramp of opaque surface brushes from `bg` to `primary`."""
key = (bg.rgb(), primary.rgb())
if self._lut_key != key:
self._lut_cache = [QBrush(_blend(bg, primary, 255.0 * (i + 0.5) / _LUT_N))
for i in range(_LUT_N)]
self._lut_key = key
return self._lut_cache
def paint(self, p: QPainter, cx: float, cy: float, r: float,
primary: QColor, accent: QColor, bg: QColor | None = None) -> None:
"""Draw the avatar with its head centre at (cx, cy).
`r` is the head's half-height in pixels — the caller owns the layout, so
the HUD can fit the head to whatever room the status line leaves it.
"""
if bg is None:
bg = QColor(0, 0, 0)
amp = self._glow
verts, norms = self._pose()
# ── aura ────────────────────────────────────────────────────────────
ar = r * 1.95
grad = QRadialGradient(cx, cy, ar)
grad.setColorAt(0.00, _c(primary, 34 + 66 * amp))
grad.setColorAt(0.38, _c(primary, 20 + 40 * amp))
grad.setColorAt(1.00, _c(primary, 0))
p.setPen(Qt.PenStyle.NoPen)
p.setBrush(QBrush(grad))
p.drawEllipse(QRectF(cx - ar, cy - ar, ar * 2, ar * 2))
# ── project ─────────────────────────────────────────────────────────
w = _CAM_D - verts[:, 2]
np.maximum(w, 0.35, out=w)
k = (_CAM_D / w) * r
xs = cx + verts[:, 0] * k
ys = cy - verts[:, 1] * k
if self.shaded:
self._paint_surface(p, xs, ys, norms, verts, primary, bg, amp)
self._paint_wire(p, xs, ys, norms, verts, primary, bg, amp)
self._paint_features(p, xs, ys, norms, r, primary, accent, bg, amp)
def _paint_surface(self, p: QPainter, xs, ys, norms, verts,
primary: QColor, bg: QColor, amp: float) -> None:
"""Fill the camera-facing triangles so the head reads as a lit volume."""
a, b, c = self._fa, self._fb, self._fc
# Flat normals, taken from each triangle's own posed geometry — NOT the
# averaged vertex normals. Averaging smears the nose, lips and brow
# relief into their neighbours and renders the face as a blank egg;
# per-facet normals are exactly what makes the anatomy visible.
fn = np.cross(verts[b] - verts[a], verts[c] - verts[a])
fn /= np.maximum(np.linalg.norm(fn, axis=1, keepdims=True), 1e-9)
# Point them outwards by agreeing with the vertex normals, which were
# oriented at build time. Flipping on the sign of n_z instead would
# negate x and y as well and scramble the lighting into moiré.
ref = norms[a] + norms[b] + norms[c]
fn *= np.sign((fn * ref).sum(1))[:, None]
nz = fn[:, 2]
area = np.abs((xs[b] - xs[a]) * (ys[c] - ys[a])
- (xs[c] - xs[a]) * (ys[b] - ys[a]))
vis = np.flatnonzero((nz > 0.015) & (area > 3.0))
if vis.size == 0:
return
fn = fn[vis]
nz = nz[vis]
ax, ay = xs[a][vis], ys[a][vis]
bx, by = xs[b][vis], ys[b][vis]
cxx, cyy = xs[c][vis], ys[c][vis]
# A rim term for the glass edge plus a key light high on the left. The
# light leans off-axis on purpose: weight it towards the camera and
# every front-facing facet returns the same value, which is a flat mask.
fres = np.clip(1.0 - nz, 0.0, 2.0) ** 1.7
lam = np.clip(fn[:, 0] * -0.55 + fn[:, 1] * 0.50 + nz * 0.52, 0.0, 1.0)
bright = 0.26 + 0.20 * fres + 0.66 * lam ** 1.05
bright *= (self._fade[a][vis] + self._fade[b][vis] + self._fade[c][vis]) / 3.0
bright *= 0.88 + 0.24 * amp
idx = np.clip((bright * _LUT_N).astype(np.int32), 0, _LUT_N - 1)
# Far facets first: the neck passes behind the jaw and the head is not
# convex around the chin.
# Sort far-to-near, but group first: neck facets all draw before head
# facets, because the two meshes interpenetrate and a pure depth sort
# interleaves them into a torn seam.
fz = (verts[a, 2][vis] + verts[b, 2][vis] + verts[c, 2][vis]) * (1.0 / 3.0)
order = np.argsort(self._fgroup[vis] * 1000.0 + fz, kind="stable")
tris = np.stack([ax, ay, bx, by, cxx, cyy], axis=1)[order].tolist()
shade = idx[order].tolist()
lut = self._lut(bg, primary)
# Aliased fills: adjacent antialiased polygons leave hairline seams, and
# the interior of a tiled surface has no silhouette worth smoothing —
# the antialiased wireframe drawn afterwards covers the outline.
p.setRenderHint(QPainter.RenderHint.Antialiasing, False)
p.setPen(Qt.PenStyle.NoPen)
for q, sh in zip(tris, shade):
p.setBrush(lut[sh])
p.drawPolygon(QPolygonF([QPointF(q[0], q[1]), QPointF(q[2], q[3]),
QPointF(q[4], q[5])]))
p.setRenderHint(QPainter.RenderHint.Antialiasing, True)
def _paint_wire(self, p: QPainter, xs, ys, norms, verts,
primary: QColor, bg: QColor, amp: float) -> None:
nz = norms[:, 2]
fres = np.abs(1.0 - np.abs(nz)) ** 1.5
if self.shaded:
# The lit surface underneath is opaque, so back-facing edges would
# float on top of the face — cull them and let the wire read as
# structure lines over skin.
front = nz > -0.05
va = np.where(front, 0.10 + 0.42 * fres, 0.0)
va += 0.30 * np.exp(-((verts[:, 1] - self._scan) / 0.13) ** 2) * front
else:
va = np.where(nz < 0.0, 0.13 + 0.26 * fres, 0.28 + 0.72 * fres)
va += 0.42 * np.exp(-((verts[:, 1] - self._scan) / 0.13) ** 2)
va *= self._fade * (0.80 + 0.45 * amp)
ea = 0.5 * (va[self._e0] + va[self._e1])
keep = np.flatnonzero(ea > _MIN_ALPHA)
if keep.size == 0:
return
# Sort by opacity bucket once so each bucket is a contiguous *slice* of
# one QLineF list; masking and rebuilding per bucket cost more than the
# drawing itself.
bucket = np.clip((ea[keep] * _BUCKETS).astype(np.int32), 0, _BUCKETS - 1)
order = np.argsort(bucket, kind="stable")
keep = keep[order]
bounds = np.searchsorted(bucket[order], np.arange(_BUCKETS + 1))
e0, e1 = self._e0[keep], self._e1[keep]
quad = np.stack([xs[e0], ys[e0], xs[e1], ys[e1]], axis=1).tolist()
lines = [QLineF(q[0], q[1], q[2], q[3]) for q in quad]
skin = _blend(bg, primary, 132) if self.shaded else bg
p.setBrush(Qt.BrushStyle.NoBrush)
for b in range(_BUCKETS):
lo, hi = int(bounds[b]), int(bounds[b + 1])
if hi <= lo:
continue
seg = lines[lo:hi]
a = 255.0 * min(1.0, (b + 0.5) / _BUCKETS)
if self.shaded:
# These sit on lit skin, so pre-mix against a representative
# *skin* tone rather than the background — same fast opaque
# path, and the lines still read as highlights over the face.
p.setPen(QPen(_blend(skin, primary, a * 0.75), 1.0))
else:
p.setPen(QPen(_blend(bg, primary, a), 1.0))
p.drawLines(seg)
# ── face ────────────────────────────────────────────────────────────────
def _ring(self, xs, ys, idx) -> QPolygonF:
return QPolygonF([QPointF(float(x), float(y))
for x, y in zip(xs[idx], ys[idx])])
def _paint_features(self, p: QPainter, xs, ys, norms, r: float,
primary: QColor, accent: QColor, bg: QColor,
amp: float) -> None:
"""Eyes, brows and the mouth cavity, drawn from the real landmark rings.
The canonical model's eyes and lips are closed skin — the geometry gives
the *shape* of the lids and mouth but no opening, so the openings are
painted here, exactly on the landmarks that bound them.
"""
face = max(0.0, math.cos(self._yaw) * math.cos(self._pitch)) ** 2
if face < 0.02:
return
lm = self._lm
vis = 1.0 - self._blink
# ── eyes ────────────────────────────────────────────────────────────
for key in ("eye_l", "eye_r"):
idx = lm[key]
ex, ey = xs[idx], ys[idx]
mid_y = float(ey.mean())
if vis < 0.999:
ey = mid_y + (ey - mid_y) * max(0.04, vis)
poly = QPolygonF([QPointF(float(a), float(b)) for a, b in zip(ex, ey)])
p.setPen(Qt.PenStyle.NoPen)
p.setBrush(QBrush(_blend(bg, primary, 22))) # socket shadow
p.drawPolygon(poly)
p.setBrush(Qt.BrushStyle.NoBrush)
p.setPen(QPen(_c(primary, 210 * face), 1.3)) # lid line
p.drawPolygon(poly)
if vis > 0.35:
br = poly.boundingRect()
gx = br.center().x() + self._gaze[0] * br.width() * 0.16
gy = br.center().y() + self._gaze[1] * br.height() * 0.20
cpt = QPointF(gx, gy)
rad = min(br.height() * 0.62, br.width() * 0.20)
p.setPen(Qt.PenStyle.NoPen)
p.setBrush(QBrush(_c(accent, (70 + 60 * amp) * face * vis)))
p.drawEllipse(cpt, rad, rad * vis) # iris
p.setBrush(QBrush(_c(accent, 245 * face * vis)))
p.drawEllipse(cpt, rad * 0.42, rad * 0.42 * vis) # pupil
# ── brows ───────────────────────────────────────────────────────────
p.setBrush(Qt.BrushStyle.NoBrush)
p.setPen(QPen(_c(primary, 150 * face), 1.7))
for key in ("brow_l", "brow_r"):
idx = lm[key]
p.drawPolyline(self._ring(xs, ys, idx))
# ── mouth ───────────────────────────────────────────────────────────
inner = self._ring(xs, ys, lm["lips_in"])
open_h = inner.boundingRect().height()
p.setPen(Qt.PenStyle.NoPen)
if self._mouth > 0.02:
# The cavity is dark but never pure black — a black oval on a glowing
# head reads as a hole, not a mouth. Tinting it with the theme keeps
# it part of the hologram.
p.setBrush(QBrush(_blend(bg, primary, 16 + 26 * self._mouth)))
p.drawPolygon(inner)
# Upper teeth: a bright strip hanging from the upper lip. It is the
# single cheapest thing that makes an open mouth look like speech.
ux, uy = xs[self._lip_up], ys[self._lip_up]
th = open_h * 0.30
pts = [QPointF(float(x), float(y)) for x, y in zip(ux, uy)]
pts += [QPointF(float(x), float(y) + th)
for x, y in zip(ux[::-1], uy[::-1])]
p.setBrush(QBrush(_blend(bg, primary, 150 + 60 * self._mouth)))
p.drawPolygon(QPolygonF(pts))
# A warm pool at the back of the throat, strongest when wide open.
p.setBrush(QBrush(_c(accent, 40 * self._mouth * face)))
p.drawPolygon(inner)
p.setBrush(Qt.BrushStyle.NoBrush)
p.setPen(QPen(_c(primary, (150 + 70 * self._mouth) * face), 1.3))
p.drawPolygon(inner) # lip edge
p.setPen(QPen(_c(primary, 110 * face), 1.1))
p.drawPolygon(self._ring(xs, ys, lm["lips_out"]))
+351
View File
@@ -0,0 +1,351 @@
"""
Human head mesh for the HUD avatar.
The face is **real measured human geometry** — MediaPipe's canonical face model
(`core/face_model.obj`, Apache-2.0, 468 vertices / 898 triangles), which carries
actual eyelids, nostrils, lips and cheekbones. Everything a formula cannot give
you comes from there.
Earlier revisions of this file generated the whole head procedurally from an
ellipsoid pushed around by gaussians. It could be tuned endlessly and still read
as an egg with a face drawn on it, because there was no human anatomy in it —
only smooth blobs. Real topology fixed in one step what parameter tweaking could
not fix at all.
What is still generated here, around that face:
* the cranium — the model is an open mask, so its 36-vertex border is swept
back and up over a skull-shaped ellipsoid and closed at the occiput;
* a tapering neck stub that fades out instead of needing shoulders;
* vertex normals, jaw-rig weights, a thinned wireframe, and the landmark
index rings (eyes, brows, lips) the renderer animates.
Coordinate system after normalisation (head-local, right-handed):
+x → viewer's right +y → up +z → out of the face
y = +1.0 crown, y = -1.0 chin, eyes land on y ≈ 0.
"""
from __future__ import annotations
import collections
from pathlib import Path
import numpy as np
_OBJ = Path(__file__).resolve().parent / "face_model.obj"
# Cranium shape, in the model's own units (chin ≈ -9.4, forehead ≈ +8.3).
# Tuned so that brow→crown is ~0.36 of the head's height, which is the real
# proportion; a taller cranium than that immediately reads as a long face even
# though the face itself is untouched measured geometry.
_SKULL_C = (0.0, 2.0, -1.0) # centre of the cranial ellipsoid
_SKULL_R = (8.4, 12.4, 8.2) # its radii
_SKULL_POLE = (0.0, 0.42, -1.0) # direction of the occiput, where the sweep closes
_SKULL_RINGS = 6
_SKULL_BLEND = 1.7 # how fast the sweep leaves the face border
_SKULL_BULGE = 1.04
_NECK_RINGS, _NECK_SEGS = 9, 14
_NECK_Z = -1.6 # the neck tube's axis, in model units
_WIRE_STRIDE = 3 # keep every n-th edge; the surface carries the form
# MediaPipe landmark rings. Verified against the geometry at build time — see
# `_check_landmarks` — so a wrong index can never silently animate the cheek.
LANDMARKS: dict[str, list[int]] = {
"eye_l": [33, 7, 163, 144, 145, 153, 154, 155, 133, 173, 157, 158, 159,
160, 161, 246],
"eye_r": [263, 249, 390, 373, 374, 380, 381, 382, 362, 398, 384, 385, 386,
387, 388, 466],
"brow_l": [70, 63, 105, 66, 107],
"brow_r": [300, 293, 334, 296, 336],
"lips_out": [61, 146, 91, 181, 84, 17, 314, 405, 321, 375, 291, 409, 270,
269, 267, 0, 37, 39, 40, 185],
"lips_in": [78, 95, 88, 178, 87, 14, 317, 402, 318, 324, 308, 415, 310,
311, 312, 13, 82, 81, 80, 191],
}
# Jaw rig, in normalised units. The pivot sits between the ears, which is where
# a real mandible hinges.
JAW_PIVOT = (0.0, 0.06, -0.34)
JAW_MAX = 0.115 # radians of drop at full amplitude (~6.6°)
# Speech barely moves a real jaw, and a talking head is watched at HUD size
# where a small, precise mouth reads better than a large one. The lip rig
# (spread / round) now carries most of the articulation, so the jaw does not
# have to swing to show that something is being said.
def _load_obj(path: Path):
verts, faces = [], []
for line in path.read_text(encoding="utf-8").splitlines():
if line.startswith("v "):
verts.append([float(x) for x in line.split()[1:4]])
elif line.startswith("f "):
faces.append([int(t.split("/")[0]) - 1 for t in line.split()[1:4]])
return np.array(verts, dtype=np.float64), np.array(faces, dtype=np.int64)
def _boundary_loop(faces: np.ndarray) -> np.ndarray:
"""Ordered ring of vertices along the open border of a triangle mesh."""
seen = collections.Counter()
for a, b, c in faces:
for e in ((a, b), (b, c), (c, a)):
seen[(min(e), max(e))] += 1
border = [e for e, n in seen.items() if n == 1]
adj = collections.defaultdict(list)
for a, b in border:
adj[a].append(b)
adj[b].append(a)
start = border[0][0]
loop, prev, cur = [start], None, start
while True:
nxt = [v for v in adj[cur] if v != prev]
if not nxt or nxt[0] == start:
break
prev, cur = cur, nxt[0]
loop.append(cur)
return np.array(loop)
def _slerp(a: np.ndarray, b: np.ndarray, t):
dot = np.clip((a * b).sum(-1, keepdims=True), -1.0, 1.0)
om = np.arccos(dot)
so = np.sin(om)
safe = np.where(so < 1e-6, 1.0, so)
out = np.where(so < 1e-6, a * (1 - t) + b * t,
(np.sin((1 - t) * om) / safe) * a + (np.sin(t * om) / safe) * b)
return out / np.maximum(np.linalg.norm(out, axis=-1, keepdims=True), 1e-9)
def _add_cranium(verts: np.ndarray, faces: np.ndarray):
"""Sweep the mask's open border back over a skull and close it at the occiput."""
loop = _boundary_loop(faces)
# Orient the loop so the generated triangles wind the same way as the face's.
centre2d = verts[loop, :2].mean(0)
ang = np.arctan2(verts[loop, 1] - centre2d[1], verts[loop, 0] - centre2d[0])
if np.diff(np.unwrap(ang)).sum() < 0:
loop = loop[::-1]
n = len(loop)
C = np.array(_SKULL_C)
R = np.array(_SKULL_R)
pole = np.array(_SKULL_POLE)
pole = pole / np.linalg.norm(pole)
chin_y = verts[:, 1].min()
rim = verts[loop] - C
rim_r = np.linalg.norm(rim, axis=1, keepdims=True)
rim_d = rim / rim_r
def ell_r(d):
return 1.0 / np.sqrt(((d / R) ** 2).sum(-1, keepdims=True))
out_v = [verts]
out_f = list(faces)
prev_idx = loop
ts = np.linspace(0.0, 1.0, _SKULL_RINGS + 1)[1:]
for t in ts:
d = _slerp(pole, rim_d, 1.0 - t)
w = (1.0 - t) ** _SKULL_BLEND # meets the rim exactly at t = 0
# A skull is fuller than the border it springs from; peak it mid-sweep.
r = ell_r(d) * (1.0 + (_SKULL_BULGE - 1.0) * np.sin(np.pi * t) ** 0.8)
ring = C + d * (w * rim_r + (1.0 - w) * r)
# Never dip below the chin: the sweep passing under the jaw would
# otherwise hang a lip of geometry below the face. Vertices that hit
# the clamp are also drawn in towards the neck axis, so the underside
# closes as a small floor instead of a flat skirt sticking out.
below = ring[:, 1] < chin_y
if below.any():
ring[below, 1] = chin_y
ring[below, 0] *= 0.55
ring[below, 2] = _NECK_Z + (ring[below, 2] - _NECK_Z) * 0.55
if t == ts[-1]:
ring = np.repeat((C + pole * ell_r(pole[None])[0])[None], n, axis=0)
base = sum(len(a) for a in out_v)
out_v.append(ring)
idx = np.arange(base, base + n)
for i in range(n):
a0, b0 = prev_idx[i], prev_idx[(i + 1) % n]
a1, b1 = idx[i], idx[(i + 1) % n]
out_f.append([a0, a1, b1])
out_f.append([a0, b1, b0])
prev_idx = idx
return np.vstack(out_v), np.array(out_f, dtype=np.int64)
def _add_neck(verts: np.ndarray, faces: np.ndarray):
"""A tapering tube dropped from inside the jaw; it fades out, so no shoulders."""
ph = np.linspace(0.0, 2.0 * np.pi, _NECK_SEGS, endpoint=False)
# Short, and flaring hard at the bottom: a straight vertical tube reads as
# a pedestal, whereas a neck that widens into the top of the shoulders
# reads as a bust — and the shorter it is, the larger the head can be drawn
# in the same HUD band.
ys = np.linspace(-5.5, -13.0, _NECK_RINGS)
d = (ys + 5.5) / -7.5
rx = 4.6 * (1.0 + 0.52 * d ** 1.9)
rz = 4.1 * (1.0 + 0.38 * d ** 1.9)
nx = rx[:, None] * np.cos(ph)[None, :]
nz = _NECK_Z + rz[:, None] * np.sin(ph)[None, :]
ny = ys[:, None] * np.ones_like(ph)[None, :]
nv = np.stack([nx.ravel(), ny.ravel(), nz.ravel()], axis=1)
base = len(verts)
idx = base + np.arange(_NECK_RINGS * _NECK_SEGS).reshape(_NECK_RINGS, _NECK_SEGS)
nf = []
for i in range(_NECK_RINGS - 1):
for j in range(_NECK_SEGS):
a, b = idx[i, j], idx[i, (j + 1) % _NECK_SEGS]
c, e = idx[i + 1, (j + 1) % _NECK_SEGS], idx[i + 1, j]
nf.append([a, b, c])
nf.append([a, c, e])
# Enough rings that the fade steps stay small. Each quad splits into one
# triangle with two top vertices and one with two bottom vertices, so a
# steep per-vertex fade gradient makes the pair land on visibly different
# brightnesses and the neck grows a sawtooth edge.
fade = np.ones(base)
nd = np.repeat(d, _NECK_SEGS)
fade = np.concatenate([fade, 1.0 - 0.72 * np.clip(nd, 0.0, 1.0) ** 1.5])
return np.vstack([verts, nv]), np.vstack([faces, np.array(nf)]), fade
def _vertex_normals(verts: np.ndarray, faces: np.ndarray,
outward: np.ndarray) -> np.ndarray:
"""Area-weighted vertex normals, flipped to agree with `outward`.
`outward` must be a per-vertex direction that genuinely points out of the
surface. A single "away from the mesh centroid" rule is NOT good enough:
down at the base of the neck that vector points almost straight down while
the real normal is horizontal, so the dot product hovers around zero and
the sign flips at random — which tears the neck into an asymmetric slab of
half-culled, half-lit triangles.
"""
a, b, c = verts[faces[:, 0]], verts[faces[:, 1]], verts[faces[:, 2]]
fn = np.cross(b - a, c - a) # length carries the area — the weighting
n = np.zeros_like(verts)
for k in range(3):
np.add.at(n, faces[:, k], fn)
n /= np.maximum(np.linalg.norm(n, axis=1, keepdims=True), 1e-9)
flip = (n * outward).sum(1) < 0
n[flip] *= -1.0
return n
def _unique_edges(faces: np.ndarray) -> np.ndarray:
e = np.vstack([faces[:, [0, 1]], faces[:, [1, 2]], faces[:, [2, 0]]])
e = np.sort(e, axis=1)
return np.unique(e, axis=0)
def _check_landmarks(verts: np.ndarray) -> None:
"""Fail loudly at build time if a landmark ring is not where it should be."""
for left, right in (("eye_l", "eye_r"), ("brow_l", "brow_r")):
cl = verts[LANDMARKS[left]].mean(0)
cr = verts[LANDMARKS[right]].mean(0)
assert cl[0] < 0 < cr[0], f"{left}/{right} are not on opposite sides"
assert abs(cl[1] - cr[1]) < 0.5, f"{left}/{right} are at different heights"
eye_y = verts[LANDMARKS["eye_l"]].mean(0)[1]
brow_y = verts[LANDMARKS["brow_l"]].mean(0)[1]
lips = verts[LANDMARKS["lips_out"]].mean(0)
assert brow_y > eye_y, "brow is not above the eye"
assert lips[1] < eye_y, "lips are not below the eyes"
assert abs(lips[0]) < 0.5, "lips are not centred"
def build_head() -> dict:
"""Assemble the full head. Called once; `get_head_mesh()` caches the result."""
verts, faces = _load_obj(_OBJ)
_check_landmarks(verts)
n_face = len(verts)
verts, faces = _add_cranium(verts, faces)
n_head = len(verts)
verts, faces, fade = _add_neck(verts, faces)
# ── normalise: crown → +1, chin → -1, eyes land on y ≈ 0 ────────────────
head_y = verts[:n_head, 1]
crown, chin = head_y.max(), head_y.min()
scale = 2.0 / (crown - chin)
centre = np.array([0.0, (crown + chin) * 0.5, 0.0])
verts = (verts - centre) * scale
# Outward reference, per part: the head is star-shaped about its own centre,
# while the neck is a tube whose outward direction is radial in x/z only.
outward = verts - np.array([0.0, verts[:n_head, 1].mean(), 0.0])
outward[n_head:] = verts[n_head:] - np.array([0.0, 0.0, _NECK_Z * scale])
outward[n_head:, 1] = 0.0
normals = _vertex_normals(verts, faces, outward)
# ── jaw rig ─────────────────────────────────────────────────────────────
# Everything below the mouth swings on the mandible, tapering to nothing at
# the ears and around the back so the nape and the neck stay put.
mouth_y = verts[LANDMARKS["lips_out"], 1].mean()
chin_y = verts[:n_head, 1].min()
jaw = np.clip((mouth_y - verts[:, 1]) / (mouth_y - chin_y), 0.0, 1.0) ** 0.8
jaw *= np.clip(0.30 + 0.85 * (verts[:, 2] / 0.55), 0.0, 1.0)
jaw[n_head:] = 0.0 # the neck never moves
jaw[LANDMARKS["lips_in"][:10]] = 1.0 # lower inner lip leads
jaw[LANDMARKS["lips_out"][:10]] = 0.95
# ── brow rig ────────────────────────────────────────────────────────────
# Raising the brows displaces the actual surface rather than sliding a drawn
# line over it, so the brow ridge relights as it lifts.
brow_y = verts[LANDMARKS["brow_l"] + LANDMARKS["brow_r"], 1].mean()
brow = np.exp(-((verts[:, 1] - brow_y) / 0.115) ** 2)
brow *= np.clip(verts[:, 2] / 0.35, 0.0, 1.0) # front of the face only
brow *= np.exp(-(verts[:, 0] / 0.42) ** 2) # fades out past the temples
brow[n_head:] = 0.0
# ── lip rig ─────────────────────────────────────────────────────────────
# Vowels are not just "how far open" — /i/ spreads the lips wide, /u/ purses
# them forward. This weight lets the renderer widen or round the mouth
# region as a whole, so the surrounding skin follows instead of tearing away
# from the landmark rings.
lip_c = verts[LANDMARKS["lips_out"]].mean(axis=0)
lips = np.exp(-((verts[:, 1] - lip_c[1]) / 0.155) ** 2)
lips *= np.exp(-(verts[:, 0] / 0.30) ** 2)
lips *= np.clip(verts[:, 2] / 0.40, 0.0, 1.0)
lips[n_head:] = 0.0
edges = _unique_edges(faces)[::_WIRE_STRIDE]
# Neck and head interpenetrate, and a painter's-algorithm sort by triangle
# depth interleaves them into a torn edge. Grouping fixes it: the neck is
# always behind the head where they overlap, so draw every neck facet first.
face_group = (faces >= n_head).all(axis=1).astype(np.int32) # 1 = neck
return {
"face_group": np.ascontiguousarray(1 - face_group, dtype=np.float32),
"brow": np.ascontiguousarray(brow, dtype=np.float32),
"lips": np.ascontiguousarray(lips, dtype=np.float32),
"lip_centre": np.ascontiguousarray(lip_c, dtype=np.float32),
"verts": np.ascontiguousarray(verts, dtype=np.float32),
"normals": np.ascontiguousarray(normals, dtype=np.float32),
"faces": np.ascontiguousarray(faces, dtype=np.int32),
"edges": np.ascontiguousarray(edges, dtype=np.int32),
"jaw": np.ascontiguousarray(jaw, dtype=np.float32),
"fade": np.ascontiguousarray(fade, dtype=np.float32),
"landmarks": {k: np.array(v, dtype=np.int32) for k, v in LANDMARKS.items()},
"n_face": n_face,
"n_head": n_head,
"span": (1.0, float(verts[:, 1].min())), # crown, bottom of the neck
}
_CACHE: dict | None = None
def get_head_mesh() -> dict:
"""Process-wide cached mesh — every HudCanvas shares the same arrays."""
global _CACHE
if _CACHE is None:
_CACHE = build_head()
return _CACHE
+161
View File
@@ -0,0 +1,161 @@
"""
core/confirm.py — a confirmation the model cannot forge.
THE PROBLEM WITH THE OLD GATE
computer_settings guarded shutdown and restart like this:
confirmed = str(params.get("confirmed", "")).lower()
if confirmed not in ("yes", "true", "1", "confirm"):
return "Please confirm by calling again with confirmed=yes."
`confirmed` is a tool parameter, which means the *model* writes it. Nothing
stops it from sending confirmed=yes on the first call, and nothing checks
that a human was ever involved. It is a convention, not a gate — and its
coverage was two actions, so deleting files and switching off the WiFi the
assistant is talking over went through with no gate at all.
THE DESIGN HERE
The confirmation token is issued by the *interface*, never by the model:
1. An action calls `request(...)` with a callable that does the real work.
2. This module hands the UI a banner with CONFIRM / CANCEL and returns
IMMEDIATELY with a sentence for the model to say out loud.
3. If — and only if — the user presses CONFIRM, the UI calls `resolve()`,
which runs the stored callable off the Qt thread.
Nothing blocks. The model keeps talking while the banner is up, so this
costs no latency at all; in fact it is cheaper than the old gate, which
burned two tool round trips (reject, then re-call) on every shutdown.
WHAT BELONGS HERE AND WHAT DOES NOT
Only genuinely irreversible things. Anything that can be reversed should be
done at once and pushed onto core/undo.py instead — undo is faster than a
question, and an assistant that asks before every action is one nobody uses.
"""
from __future__ import annotations
import threading
import time
from dataclasses import dataclass
from typing import Callable, Optional
# A pending confirmation is abandoned after this long. Chosen to outlast a
# normal "hang on, let me look at the screen" pause without leaving a live
# shutdown button sitting on the HUD for the rest of the day.
TIMEOUT_SECONDS = 90.0
@dataclass
class _Pending:
key: str
title: str
detail: str
run: Callable[[], str]
at: float
_pending: Optional[_Pending] = None
_lock = threading.Lock()
# Set once at startup by main.py. Signature: (title, detail) -> None for show,
# and () -> None for hide. Both are marshalled onto the Qt thread by the UI.
_show_cb: Optional[Callable[[str, str], None]] = None
_hide_cb: Optional[Callable[[], None]] = None
_log_cb: Optional[Callable[[str], None]] = None
def bind(show, hide, log=None) -> None:
"""Wire this module to the HUD. Called once from main.py at startup."""
global _show_cb, _hide_cb, _log_cb
_show_cb, _hide_cb, _log_cb = show, hide, log
def _log(msg: str) -> None:
if _log_cb:
try:
_log_cb(msg)
except Exception:
pass
def request(key: str, title: str, detail: str, run: Callable[[], str]) -> str:
"""Park an irreversible action behind the on-screen gate.
Returns the sentence the tool should hand back to the model — phrased as an
instruction so the assistant asks the user out loud in their own language,
rather than reading an English string verbatim."""
global _pending
if _show_cb is None:
# No interface bound (headless, or a very early call). Refuse rather
# than silently performing something irreversible.
return (f"I cannot confirm '{title}' right now because the interface is "
f"not available, so I have not done it.")
with _lock:
_pending = _Pending(key=key, title=title, detail=detail,
run=run, at=time.monotonic())
try:
_show_cb(title, detail)
except Exception as e:
with _lock:
_pending = None
return f"Could not ask for confirmation: {e}. Nothing was done."
_log(f"SYS: Awaiting confirmation — {title}")
return (
f"[CONFIRMATION_PENDING] I have put a confirmation on screen for: {title}. "
f"Say ONE short sentence in the user's own language telling them you need "
f"them to confirm it on the HUD before you do it. Do not claim it is done."
)
def resolve(accepted: bool) -> None:
"""Called by the UI when the user presses CONFIRM or CANCEL.
Runs the stored callable on a worker thread — this is invoked from the Qt
thread, and shutting the machine down from inside a button handler would
freeze the interface on its way out."""
global _pending
with _lock:
p, _pending = _pending, None
if _hide_cb:
try:
_hide_cb()
except Exception:
pass
if p is None:
return
if time.monotonic() - p.at > TIMEOUT_SECONDS:
_log(f"SYS: Confirmation expired — {p.title}")
return
if not accepted:
_log(f"SYS: Cancelled — {p.title}")
return
def _worker():
try:
result = p.run() or "Done."
_log(f"SYS: Confirmed — {p.title}. {result}")
except Exception as e:
_log(f"ERR: {p.title} failed — {e}")
threading.Thread(target=_worker, daemon=True,
name=f"confirm-{p.key}").start()
def pending_title() -> str:
"""'' when nothing is waiting. Lets an action avoid stacking two banners."""
with _lock:
if _pending is None:
return ""
if time.monotonic() - _pending.at > TIMEOUT_SECONDS:
return ""
return _pending.title
+1378
View File
File diff suppressed because it is too large. Load diff
+173
View File
@@ -0,0 +1,173 @@
"""
Push-to-talk — hold a key, speak, release.
Why this exists
---------------
Wake-word is hands-free but it is not always what you want: in a meeting, in a
noisy room, or when you simply do not feel like saying a name out loud, a key
you hold is faster and never mishears. It is also the natural way to use the
assistant while another window has focus.
How it works, and what it costs
-------------------------------
Zero new dependencies, best available mechanism per platform:
* **Windows** — `GetAsyncKeyState` polled from one small thread. This is
deliberately *not* `RegisterHotKey`, which only reports a press: push-to-talk
needs the release too, and it needs to work while another application has
focus. Polling two virtual-key codes 30 times a second is a rounding error of
CPU and needs no message loop.
* **macOS / Linux** — no portable way to read global key state without pulling
in a new package or asking for accessibility permissions, so the chord is
bound as an application shortcut instead: it works whenever the assistant's
window has focus. `scope` reports which of the two you got, so the UI can say
so honestly rather than pretending.
The class never raises. If the platform hook cannot be installed it simply
reports `scope == "window"` and the Qt shortcut carries it.
"""
from __future__ import annotations
import platform
import threading
import time
from typing import Callable
_OS = platform.system()
# The default chord. Ctrl+Space is free in most desktop environments and is the
# same finger shape on every keyboard layout, which matters for a worldwide app.
DEFAULT_CHORD = ("ctrl", "space")
# Windows virtual-key codes for the names we accept.
_VK = {
"ctrl": 0x11, "shift": 0x10, "alt": 0x12,
"space": 0x20, "f8": 0x77, "f9": 0x78, "f10": 0x79,
"capslock": 0x14, "insert": 0x2D,
}
# Qt key sequence text for the same chord, used by the windowed fallback.
_QT_NAME = {"ctrl": "Ctrl", "shift": "Shift", "alt": "Alt", "space": "Space",
"f8": "F8", "f9": "F9", "f10": "F10",
"capslock": "CapsLock", "insert": "Ins"}
_POLL_HZ = 30.0
# A key has to be down this long before we call it speech. It stops a stray
# brush of the chord from opening the microphone.
_DEBOUNCE_S = 0.06
def chord_label(chord=DEFAULT_CHORD) -> str:
"""Human-readable name of the chord, for the UI and the logs."""
return "+".join(_QT_NAME.get(k, k.title()) for k in chord)
def qt_sequence(chord=DEFAULT_CHORD) -> str:
"""The same chord as a QKeySequence string."""
return "+".join(_QT_NAME.get(k, k.title()) for k in chord)
class PushToTalk:
"""Calls `on_change(held: bool)` whenever the chord is pressed or released.
Start it once; it is safe to start and stop repeatedly, and safe to stop a
detector that never started.
"""
def __init__(self, on_change: Callable[[bool], None], chord=DEFAULT_CHORD):
self._on_change = on_change
self._chord = tuple(chord)
self._thread: threading.Thread | None = None
self._stop = threading.Event()
self._held = False
self._scope = "window"
# ── state ───────────────────────────────────────────────────────────────
@property
def held(self) -> bool:
return self._held
@property
def scope(self) -> str:
"""'global' once a system-wide hook is running, else 'window'."""
return self._scope
@property
def label(self) -> str:
return chord_label(self._chord)
# ── lifecycle ───────────────────────────────────────────────────────────
def start(self) -> str:
"""Begin watching. Returns the scope actually achieved."""
self.stop()
self._stop.clear()
if _OS == "Windows" and self._can_poll():
self._scope = "global"
self._thread = threading.Thread(
target=self._poll_loop, name="push-to-talk", daemon=True)
self._thread.start()
else:
self._scope = "window"
return self._scope
def stop(self) -> None:
self._stop.set()
t, self._thread = self._thread, None
if t is not None and t.is_alive():
t.join(timeout=1.0)
self._set_held(False)
# ── the windowed fallback drives this directly ──────────────────────────
def set_held(self, held: bool) -> None:
"""Feed a press/release from a Qt shortcut (non-Windows, or no hook)."""
self._set_held(bool(held))
# ── internals ───────────────────────────────────────────────────────────
def _can_poll(self) -> bool:
try:
import ctypes
ctypes.windll.user32.GetAsyncKeyState # noqa: B018 — presence check
return all(k in _VK for k in self._chord)
except Exception:
return False
def _set_held(self, held: bool) -> None:
if held == self._held:
return
self._held = held
try:
self._on_change(held)
except Exception:
pass # a listener fault must never kill the watcher
def _poll_loop(self) -> None:
import ctypes
user32 = ctypes.windll.user32
codes = [_VK[k] for k in self._chord]
period = 1.0 / _POLL_HZ
down_since = 0.0
while not self._stop.is_set():
try:
# The high bit of the return value is "currently down".
down = all(user32.GetAsyncKeyState(c) & 0x8000 for c in codes)
except Exception:
break # driver or session teardown — fall back to windowed
now = time.monotonic()
if down:
if down_since == 0.0:
down_since = now
elif now - down_since >= _DEBOUNCE_S:
self._set_held(True)
else:
down_since = 0.0
self._set_held(False)
self._stop.wait(period)
self._set_held(False)
self._scope = "window"
+275
View File
@@ -0,0 +1,275 @@
"""
Text → mouth shape, fused with the audio the avatar is actually speaking.
Why both sources
----------------
Formant analysis of the audio (see `_pcm_visemes` in main.py) gives excellent
*timing* and a decent read on vowels, but it is blind to exactly the consonants
lip-reading depends on. /m/, /b/ and /p/ are made with the lips pressed shut,
and nothing in the spectrum reliably says "the lips are closed" — a nasal /m/
and a nasal /n/ look nearly identical to a filter bank while looking completely
different on a face.
The transcript knows those consonants for certain. So the text supplies *which
shape*, the audio supplies *when* and *how strongly*, and the two are blended.
If the transcript is late or missing the mouth silently falls back to the
audio-only shape, which is what the previous version did on its own.
Language independence
---------------------
There is no per-language table here. Every character is reduced to one of the
26 bare Latin letters — by Unicode decomposition for accents, by transliteration
for Cyrillic and Greek — and articulation is looked up on that. So Turkish,
English, German, French, Spanish, Polish, Vietnamese, Russian, Ukrainian and
Greek all work from the same twenty-odd rules, and adding a language costs
nothing because there is nothing to add.
Scripts whose spelling does not reveal pronunciation (CJK, Arabic, Devanagari,
Hebrew, Thai) are detected by coverage and skipped, and the mouth runs on the
audio-only shape — which is itself language-independent, being physics. The
result is never wrong, only less detailed.
"""
from __future__ import annotations
import unicodedata
from collections import deque
# (openness 0..1, width -1..+1, closure 0..1)
# closure forces the lips together regardless of loudness — it is the whole
# reason the transcript is worth consulting.
VISEMES: dict[str, tuple[float, float, float]] = {
"REST": (0.00, 0.00, 0.00),
"AA": (0.92, -0.05, 0.00), # a
"E": (0.52, 0.42, 0.00), # e
"I": (0.20, 0.62, 0.00), # i, ı
"O": (0.55, -0.52, 0.00), # o, ö
"U": (0.26, -0.74, 0.00), # u, ü, w
"MBP": (0.00, 0.00, 1.00), # m, b, p — lips pressed shut
"FV": (0.10, 0.22, 0.55), # f, v — lower lip to the teeth
"S": (0.16, 0.42, 0.00), # s, ş, z, c, ç, j
"L": (0.36, 0.18, 0.00), # l
"TD": (0.28, 0.12, 0.00), # t, d, n
"K": (0.30, -0.04, 0.00), # k, g, ğ, h
"R": (0.28, -0.16, 0.00), # r
}
# Relative duration of each class. Vowels carry the syllable; plosives are a tap.
_DUR = {"REST": 1.0, "AA": 1.15, "E": 1.05, "I": 1.0, "O": 1.1, "U": 1.05,
"MBP": 0.5, "FV": 0.8, "S": 0.9, "L": 0.65, "TD": 0.5, "K": 0.55,
"R": 0.55}
# Articulation is a property of the *sound*, not of a language, so the table is
# keyed on the 26 bare Latin letters and every script reaches it by reduction:
# * diacritics are stripped by Unicode decomposition (é→e, ü→u, ế→e, ł→l …),
# which covers every Latin-script language at once rather than one at a time;
# * letters that do not decompose get a short explicit entry below;
# * Cyrillic and Greek transliterate into the same 26 letters.
# Anything else — CJK, Arabic, Devanagari, Hebrew, Thai — is not derivable from
# its written form without a pronunciation dictionary, so those simply fall back
# to the audio-only mouth. That is a clean degradation, not a missing language.
_LETTER = {
"a": "AA",
"e": "E",
"i": "I", "y": "I",
"o": "O",
"u": "U", "w": "U",
"b": "MBP", "p": "MBP", "m": "MBP",
"f": "FV", "v": "FV",
"s": "S", "z": "S", "c": "S", "j": "S", "x": "S",
"l": "L",
"t": "TD", "d": "TD", "n": "TD",
"k": "K", "g": "K", "h": "K", "q": "K",
"r": "R",
}
# Letters with no Unicode decomposition into a Latin base.
_UNDECOMPOSED = {
"ı": "i", "ø": "o", "đ": "d", "ħ": "h", "ŀ": "l", "ŧ": "t",
"ß": "s", "æ": "a", "œ": "o", "þ": "t", "ð": "d", "ŋ": "n",
"ł": "l",
}
_CYRILLIC = {
"а": "a", "б": "b", "в": "v", "г": "g", "д": "d", "е": "e", "ё": "e",
"ж": "j", "з": "z", "и": "i", "й": "i", "к": "k", "л": "l", "м": "m",
"н": "n", "о": "o", "п": "p", "р": "r", "с": "s", "т": "t", "у": "u",
"ф": "f", "х": "h", "ц": "s", "ч": "s", "ш": "s", "щ": "s", "ъ": "",
"ы": "i", "ь": "", "э": "e", "ю": "u", "я": "a",
"і": "i", "ї": "i", "є": "e", "ґ": "g", "ў": "u",
}
_GREEK = {
"α": "a", "β": "v", "γ": "g", "δ": "d", "ε": "e", "ζ": "z", "η": "i",
"θ": "t", "ι": "i", "κ": "k", "λ": "l", "μ": "m", "ν": "n", "ξ": "s",
"ο": "o", "π": "p", "ρ": "r", "σ": "s", "ς": "s", "τ": "t", "υ": "i",
"φ": "f", "χ": "h", "ψ": "s", "ω": "o",
}
# Below this share of mappable letters the text is in a script we cannot read
# phonetically, and forcing shapes onto it would be worse than not trying.
_MIN_COVERAGE = 0.55
# English spellings that do not survive letter-by-letter reading.
_DIGRAPH = {
"sh": "S", "ch": "S", "ts": "S",
"th": "TD", "ck": "K", "ng": "K", "gh": "K",
"ph": "FV",
"oo": "U", "ou": "O", "ow": "O", "wh": "U",
"ee": "I", "ea": "I", "ie": "I",
"qu": "K",
}
_PAUSE = set(".,;:!?…\n")
def to_latin(ch: str) -> str:
"""Reduce any character to a bare Latin letter, or "" if it has none.
This is what makes the mouth language-agnostic: one reduction step replaces
a per-language spelling table.
"""
c = ch.lower()
if "a" <= c <= "z":
return c
if c in _UNDECOMPOSED:
return _UNDECOMPOSED[c]
if c in _CYRILLIC:
return _CYRILLIC[c]
if c in _GREEK:
return _GREEK[c]
# Strip combining marks: é→e, ü→u, ş→s, ğ→g, ế→e, ñ→n, å→a …
base = "".join(k for k in unicodedata.normalize("NFD", c)
if not unicodedata.combining(k))
if len(base) == 1 and "a" <= base <= "z":
return base
if base and base != c: # e.g. fi → fi, take the first
return to_latin(base[0])
return ""
def coverage(text: str) -> float:
"""Fraction of the letters in `text` we can reduce to a Latin sound."""
letters = [c for c in (text or "") if c.isalpha()]
if not letters:
return 0.0
return sum(1 for c in letters if to_latin(c)) / len(letters)
def text_to_visemes(text: str) -> list[tuple[str, float]]:
"""Split a line of speech into (viseme, duration-weight) pairs.
Returns [] for scripts whose written form does not reveal pronunciation, so
the caller falls back to the audio-only mouth instead of miming nonsense.
"""
s = (text or "").lower()
if coverage(s) < _MIN_COVERAGE:
return []
out: list[tuple[str, float]] = []
i, n = 0, len(s)
while i < n:
ch = s[i]
if ch in _PAUSE:
out.append(("REST", 1.4))
i += 1
continue
if ch.isspace():
# A word gap is a beat, not a closed mouth — closing between every
# word makes the avatar look like it is chewing.
if out and out[-1][0] != "REST":
out.append((out[-1][0], 0.35))
i += 1
continue
# Digraphs are an orthographic quirk of Latin spelling; check them on
# the reduced letters so "SCH"/"Sch" and accented forms match too.
two = to_latin(ch) + (to_latin(s[i + 1]) if i + 1 < n else "")
if len(two) == 2 and two in _DIGRAPH:
v = _DIGRAPH[two]
i += 2
else:
base = to_latin(ch)
i += 1
if not base:
continue
v = _LETTER.get(base)
if v is None:
continue
# A doubled letter is one sound in every orthography we handle here.
if out and out[-1][0] == v:
continue
out.append((v, _DUR[v]))
return out
class VisemeStream:
"""Fuses the transcript's shape sequence onto the audio's timing.
Thread note: `feed_text` runs on the receive coroutine and `frames` on the
playback coroutine. Both live in the same asyncio loop, and `deque` append
and popleft are atomic, so no lock is needed.
"""
# Seconds a phoneme occupies at a normal speaking rate. The clock adapts
# between these when the queue runs long (the model is talking fast) or
# short (it is trailing off).
_MIN_STEP = 0.045
_MAX_STEP = 0.105
def __init__(self) -> None:
self._q: deque[tuple[str, float]] = deque()
self._cur = ("REST", 1.0)
self._carry = 0.0
def reset(self) -> None:
self._q.clear()
self._cur = ("REST", 1.0)
self._carry = 0.0
def feed_text(self, text: str) -> None:
for item in text_to_visemes(text):
self._q.append(item)
# Never let a stalled turn pile up an unbounded backlog.
while len(self._q) > 600:
self._q.popleft()
@property
def pending(self) -> int:
return len(self._q)
def _step_seconds(self) -> float:
# A long backlog means speech is outrunning the clock; shorten the step
# so the mouth catches up instead of drifting further behind the voice.
backlog = min(1.0, len(self._q) / 45.0)
return self._MAX_STEP - (self._MAX_STEP - self._MIN_STEP) * backlog
def frames(self, audio, hop: float):
"""Blend audio frames [(level, openness, width)] with the text queue."""
out = []
for level, a_open, a_wide in audio:
if level <= 0.0:
# Silence: let the queue wait rather than burning through it
# during a pause, or the mouth ends up ahead of the voice.
out.append((0.0, 0.0, 0.0))
continue
self._carry += hop / max(1e-3, self._step_seconds() * self._cur[1])
while self._carry >= 1.0 and self._q:
self._cur = self._q.popleft()
self._carry -= 1.0
if self._carry >= 1.0:
self._carry = 1.0 # queue empty — hold the last shape
t_open, t_wide, closure = VISEMES.get(self._cur[0], VISEMES["REST"])
if self._q or self._cur[0] != "REST":
# Text leads the shape; the audio keeps it honest so a bad
# transcript alignment still tracks the real voice.
o = 0.72 * t_open + 0.28 * a_open
w = 0.78 * t_wide + 0.22 * a_wide
else:
o, w = a_open, a_wide
closure = 0.0
o *= 1.0 - closure
out.append((level, max(0.0, min(1.0, o)), max(-1.0, min(1.0, w))))
return out
+211
View File
@@ -0,0 +1,211 @@
"""
Local wake-word detection for JARVIS ("Hey Jarvis").
Design goals:
• ZERO cost when the feature is off — openwakeword is imported ONLY inside
start()/install helpers, never at module load. If the user never enables
wake word, none of this touches the app.
• ZERO latency on the audio path — the microphone callback only ever does a
cheap, non-blocking queue push (feed()); the actual model inference runs in
this module's own background thread, so the real-time audio thread and the
Gemini stream are never slowed.
• Fully local & offline — audio fed here never leaves the machine; there is no
network call except the one-time model download the user triggers from the UI.
openwakeword ships small ONNX models (a few MB each) and runs comfortably on a
CPU. The pretrained wake phrase used here is "Hey Jarvis".
"""
from __future__ import annotations
import queue
import subprocess
import sys
import threading
from pathlib import Path
from typing import Callable
# Pretrained openwakeword model that listens for "Hey Jarvis".
WAKE_MODEL = "hey_jarvis"
# Score in [0,1]; above this counts as a detection. Tunable per environment.
DEFAULT_THRESHOLD = 0.5
# Mic frames arrive at 16 kHz int16; this is just the detector's input rate.
SAMPLE_RATE = 16000
def is_installed() -> bool:
"""True if the openwakeword package is importable (no model check)."""
try:
import importlib.util
return importlib.util.find_spec("openwakeword") is not None
except Exception:
return False
def is_ready() -> bool:
"""True if openwakeword is installed AND its model files are present on disk.
This is a cheap, DETERMINISTIC file-existence check. It deliberately does NOT
construct a Model to probe readiness — doing that is slow and, worse, can clash
with the detector's own Model when it's already running, which intermittently
returned False and made the UI flicker to 'not downloaded'. Never raises.
"""
if not is_installed():
return False
try:
import openwakeword
models_dir = Path(openwakeword.__file__).resolve().parent / "resources" / "models"
if not models_dir.is_dir():
return False
has_wake = (any(models_dir.glob(f"{WAKE_MODEL}*.onnx"))
or any(models_dir.glob(f"{WAKE_MODEL}*.tflite")))
has_mel = (any(models_dir.glob("melspectrogram*.onnx"))
or any(models_dir.glob("melspectrogram*.tflite")))
has_emb = (any(models_dir.glob("embedding_model*.onnx"))
or any(models_dir.glob("embedding_model*.tflite")))
return bool(has_wake and has_mel and has_emb)
except Exception:
return False
def install_and_download(logger: Callable[[str], None] = print,
notify: Callable[[str], None] | None = None) -> tuple[bool, str]:
"""
One-click setup for the UI button: pip-install openwakeword if missing, then
download the wake model. Returns (ok, message). Never raises — every failure
is reported through the returned message and the logger.
"""
_tell = notify or (lambda _msg: None)
try:
if not is_installed():
logger("Wake word: installing openwakeword (one-time)…")
_tell("Wake word: installing openwakeword (one-time)…")
r = subprocess.run(
[sys.executable, "-m", "pip", "install", "openwakeword"],
capture_output=True, text=True,
)
if r.returncode != 0:
tail = (r.stderr or r.stdout or "").strip().splitlines()[-1:] or [""]
return False, f"pip install failed: {tail[0][:160]}"
# Download the pretrained melspectrogram/embedding + wake models.
logger("Wake word: downloading models…")
_tell("Wake word: downloading models…")
try:
import openwakeword.utils as _u
try:
_u.download_models([WAKE_MODEL])
except TypeError:
_u.download_models() # older signature downloads the default set
except Exception as e:
return False, f"model download failed: {e}"
if not is_ready():
return False, "installed, but the wake model could not be loaded."
logger("Wake word: ready.")
return True, "Wake word installed and ready."
except Exception as e:
return False, f"setup error: {e}"
class WakeWordDetector:
"""
Runs the wake model in a dedicated thread. The mic thread calls feed() with
raw int16 frames; detections invoke on_detect() (called from this thread —
the callback must marshal to whatever loop/UI it needs).
"""
def __init__(self, on_detect: Callable[[], None],
threshold: float = DEFAULT_THRESHOLD,
logger: Callable[[str], None] = print,
notify: Callable[[str], None] | None = None):
self._on_detect = on_detect
self._threshold = threshold
self._logger = logger
# See PluginRegistry: `logger` is the console and gets everything,
# `notify` is the activity log and gets only what the user must act on.
self._notify = notify or (lambda _msg: None)
self._queue: queue.Queue = queue.Queue(maxsize=50)
self._thread: threading.Thread | None = None
self._running = False
self._model = None
self._ready = False
def start(self) -> bool:
"""Load the model and spawn the inference thread. Returns True on success.
Safe to call again — a no-op if already running. Never raises."""
if self._running:
return True
try:
from openwakeword.model import Model
self._model = Model(wakeword_models=[WAKE_MODEL], inference_framework="onnx")
except Exception as e:
self._logger(f"Wake word: could not load model — {e}")
self._notify("Wake word unavailable — use the WAKE NOW button.")
self._model = None
return False
self._running = True
self._ready = True
self._thread = threading.Thread(target=self._loop, daemon=True, name="WakeWordThread")
self._thread.start()
self._logger("Wake word: listening for 'Hey Jarvis'.")
return True
def stop(self) -> None:
self._running = False
# unblock the thread if it's waiting on the queue
try:
self._queue.put_nowait(None)
except Exception:
pass
self._model = None
self._ready = False
@property
def ready(self) -> bool:
return self._ready
def feed(self, frame_int16) -> None:
"""Called from the mic callback (real-time thread). Must stay cheap and
never block — the frame is copied and dropped if the queue is backed up."""
if not self._running:
return
try:
# frame_int16 is a numpy int16 array (possibly 2-D mono) — flatten to 1-D
data = frame_int16[:, 0].copy() if getattr(frame_int16, "ndim", 1) > 1 else frame_int16.copy()
self._queue.put_nowait(data)
except queue.Full:
pass
except Exception:
pass
def _loop(self) -> None:
import numpy as np
while self._running:
try:
frame = self._queue.get()
if frame is None or not self._running:
break
scores = self._model.predict(np.asarray(frame, dtype=np.int16))
score = 0.0
if isinstance(scores, dict):
# match the jarvis model regardless of exact key suffix
for k, v in scores.items():
if "jarvis" in k.lower():
score = max(score, float(v))
if score == 0.0 and scores:
score = max(float(v) for v in scores.values())
if score >= self._threshold:
# drain any backlog so we don't double-fire on the same utterance
self._drain()
try:
self._on_detect()
except Exception as e:
self._logger(f"Wake word: on_detect error — {e}")
except Exception as e:
self._logger(f"Wake word: inference error — {e}")
def _drain(self) -> None:
try:
while True:
self._queue.get_nowait()
except Exception:
pass
+94
View File
@@ -0,0 +1,94 @@
"""AvatarPy: assistente personale con volto animato (interfaccia da Mark-LIV, motori propri)."""
from __future__ import annotations
import json
import platform
import sys
from pathlib import Path
BASE_DIR = Path(__file__).resolve().parent
sys.path.insert(0, str(BASE_DIR))
from avatar.settings import CONFIG_DIR, Settings # noqa: E402
def ensure_identity_file() -> None:
"""L'interfaccia legge config/api_keys.json per nome, colore e stato di configurazione."""
CONFIG_DIR.mkdir(parents=True, exist_ok=True)
f = CONFIG_DIR / "api_keys.json"
data: dict = {}
if f.exists():
try:
data = json.loads(f.read_text(encoding="utf-8"))
except Exception:
data = {}
changed = False
if not data.get("os_system"):
data["os_system"] = {"Darwin": "mac", "Windows": "windows"}.get(platform.system(), "linux")
changed = True
if not data.get("assistant_name"):
data["assistant_name"] = "Ava"
changed = True
if changed:
f.write_text(json.dumps(data, indent=4, ensure_ascii=False), encoding="utf-8")
def main() -> None:
ensure_identity_file()
settings = Settings()
try:
from PyQt6 import QtWebEngineWidgets # noqa: F401 (va importato prima della QApplication)
except Exception as err:
print(f"[avatar3d] WebEngine non disponibile: {err}")
from ui import JarvisUI # crea la QApplication
from avatar.assistant import Assistant
from avatar.settings_dialog import SettingsDialog
ui = JarvisUI(str(BASE_DIR / "core" / "face_model.obj"))
ui._win._log._ai_name_lc = ui.assistant_name.lower()
assistant = Assistant(ui, settings)
ui.on_text_command = assistant.handle_text
ui.on_interrupt = assistant.interrupt
ui.on_voice_change = lambda: assistant.reload_voice(from_customise=True)
ui.on_audio_device_change = assistant.reopen_audio
ui.on_push_to_talk = assistant.set_push_to_talk
ui.ptt_hold = assistant.ptt_hold
ui.on_wake_toggle = assistant.on_wake_toggle
ui.on_wake_manual = assistant.on_wake_manual
ui.wake_get_state = assistant.wake_get_state
from avatar.plugins import registry
from core import confirm
confirm.bind(ui.show_confirm, ui.hide_confirm, ui.write_log)
registry.ctx = {"say": assistant.say, "log": ui.write_log, "confirm": confirm.request, "player": ui}
ui.get_plugins = registry.list_for_ui
ui.get_plugin_settings = lambda: []
ui.request_say = assistant.say
# Questi due sono attributi della finestra, non della facciata JarvisUI.
ui._win.on_new_conversation = assistant.reset_conversation
def open_settings() -> None:
SettingsDialog(settings, assistant.reconfigure, parent=ui._win).exec()
ui._win.on_engine_settings = open_settings
if "--smoke" in sys.argv:
from PyQt6.QtCore import QTimer
QTimer.singleShot(3000, lambda: (print("SMOKE OK"), ui._app.quit()))
else:
if not settings.ready:
from PyQt6.QtCore import QTimer
QTimer.singleShot(800, open_settings)
assistant.start()
try:
ui.root.mainloop()
finally:
try:
from avatar import whatsapp_bridge as wb
wb.stop()
except Exception:
pass
if __name__ == "__main__":
main()
View File
Whitespace-only changes.
+384
View File
@@ -0,0 +1,384 @@
import json
import sys
from pathlib import Path
def get_base_dir() -> Path:
if getattr(sys, "frozen", False):
return Path(sys.executable).parent
return Path(__file__).resolve().parent.parent
BASE_DIR = get_base_dir()
CONFIG_DIR = BASE_DIR / "config"
CONFIG_FILE = CONFIG_DIR / "api_keys.json"
def ensure_config_dir() -> None:
CONFIG_DIR.mkdir(parents=True, exist_ok=True)
def config_exists() -> bool:
return CONFIG_FILE.exists()
def save_api_keys(gemini_api_key: str) -> None:
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
data["gemini_api_key"] = gemini_api_key.strip()
CONFIG_FILE.write_text(
json.dumps(data, indent=2),
encoding="utf-8"
)
def load_api_keys() -> dict:
if not CONFIG_FILE.exists():
return {}
try:
return json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception as e:
print(f"❌ Failed to load api_keys.json: {e}")
return {}
def get_gemini_key() -> str | None:
return load_api_keys().get("gemini_api_key")
def is_configured() -> bool:
key = get_gemini_key()
return bool(key and len(key) > 15)
def get_assistant_name() -> str:
"""Return the configured assistant name, or 'JARVIS' if not set."""
return load_api_keys().get("assistant_name", "JARVIS") or "JARVIS"
def get_user_name() -> str:
"""Return the configured user name for addressing."""
return load_api_keys().get("user_name", "")
def save_assistant_config(assistant_name: str, user_name: str) -> None:
"""Persist assistant name and user name to config."""
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
data["assistant_name"] = assistant_name.strip() or "JARVIS"
data["user_name"] = user_name.strip()
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
# ── Assistant voice ──────────────────────────────────────────────────────────
# Gemini Live prebuilt voices. Names are proper nouns — identical in every
# language, so this list is safe to show verbatim in any locale.
AVAILABLE_VOICES = ["Sara", "Nicola", "Sistema"] # AvatarPy: Kokoro (Sara, Nicola) o voce di macOS
DEFAULT_VOICE = "Sara"
def get_voice() -> str:
"""Return the configured Live voice, falling back to the default if unset
or if the stored value is not a voice we recognise."""
v = load_api_keys().get("voice_name", DEFAULT_VOICE) or DEFAULT_VOICE
return v if v in AVAILABLE_VOICES else DEFAULT_VOICE
def save_voice(voice_name: str) -> None:
"""Persist the chosen Live voice. Unknown names collapse to the default so a
bad value can never reach the API and break the session."""
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
v = (voice_name or "").strip()
data["voice_name"] = v if v in AVAILABLE_VOICES else DEFAULT_VOICE
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
def get_wake_word_enabled() -> bool:
"""Whether local wake-word gating is on (assistant sleeps until 'Hey Jarvis')."""
return load_api_keys().get("wake_word_enabled", False)
def save_wake_word_enabled(enabled: bool) -> None:
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
data["wake_word_enabled"] = bool(enabled)
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
def get_push_to_talk_enabled() -> bool:
"""Hold-a-key-to-speak. When on, the mic is closed unless the chord is held."""
return load_api_keys().get("push_to_talk_enabled", False)
def save_push_to_talk_enabled(enabled: bool) -> None:
_save_flag("push_to_talk_enabled", enabled)
HUD_STYLES = ("face", "core", "3d") # AvatarPy: anche l'avatar 3D
def get_hud_style() -> str:
"""Which centrepiece the HUD draws: the animated head, or the reactor core.
Taste, not capability — both render in the same software painter and cost
about the same. Defaults to the head because that is what MARK LIV shipped
with; anyone who preferred the older look can switch back in ⚙ and the
choice survives a restart.
"""
v = str(load_api_keys().get("hud_style", "face")).strip().lower()
return v if v in HUD_STYLES else "face"
def save_hud_style(style: str) -> None:
s = str(style or "").strip().lower()
_save_flag("hud_style", s if s in HUD_STYLES else "face")
# ── Live-session tuning ──────────────────────────────────────────────────────
# Everything here is optional and has a working default, so an untouched
# config behaves exactly like a configured one. Each value is also a way out:
# if a future model dislikes one of these, set it back and nothing else changes.
def get_thinking_enabled() -> bool:
"""Whether the Live model may spend tokens thinking before it answers.
Off by default. A voice assistant is judged on how fast it starts talking,
and the reasoning that actually needs deliberation in this app is delegated
to the planning tools, which run on a separate non-Live model.
"""
return bool(load_api_keys().get("thinking_enabled", False))
def save_thinking_enabled(enabled: bool) -> None:
_save_flag("thinking_enabled", enabled)
def get_turn_tuning() -> dict:
"""How eagerly the server decides you have stopped speaking.
OFF by default, and that default was earned. Cutting turns shorter looks
like a free speed win and is not: proactive audio has to judge whether an
utterance was even addressed to the assistant, and a turn clipped early
gives it less to judge, so it stays quiet — and the reply to your first
sentence only arrives once your second one has given it enough context.
That reads as the assistant being a turn behind, which is far worse than
the fraction of a second the tuning saves.
Turn it on with "turn_tuning": {"enabled": true} if your own microphone and
speaking pace suit it. `silence_ms` is the one that is felt: the pause the
server sits through before accepting your turn is over.
"""
cfg = load_api_keys().get("turn_tuning")
cfg = cfg if isinstance(cfg, dict) else {}
def _int(key, default, lo, hi):
try:
return max(lo, min(hi, int(cfg.get(key, default))))
except (TypeError, ValueError):
return default
return {
"enabled": bool(cfg.get("enabled", False)),
"silence_ms": _int("silence_ms", 550, 200, 3000),
"prefix_ms": _int("prefix_ms", 150, 0, 1000),
# "high" = quicker to decide speech has ended.
"end_sensitivity": str(cfg.get("end_sensitivity", "high")).lower(),
"start_sensitivity": str(cfg.get("start_sensitivity", "default")).lower(),
}
def save_turn_tuning(values: dict) -> None:
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
cur = data.get("turn_tuning")
cur = dict(cur) if isinstance(cur, dict) else {}
cur.update(values or {})
data["turn_tuning"] = cur
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
def get_proactive_audio_enabled() -> bool:
"""Whether the model gets to decide an utterance was not aimed at it and
stay quiet.
On by default — it is what stops the assistant answering the room. But it
is also the first thing to switch off if replies ever seem to arrive a turn
late: what looks like lag is usually the model having judged your previous
sentence as not addressed to it, and only changing its mind once the next
one arrives.
"""
return bool(load_api_keys().get("proactive_audio", True))
def save_proactive_audio_enabled(enabled: bool) -> None:
_save_flag("proactive_audio", enabled)
MEDIA_RESOLUTIONS = ("default", "low", "medium", "high")
def get_media_resolution() -> str:
"""How finely the model tokenises the screenshots and camera frames it is
sent. 'medium' keeps on-screen text readable at a fraction of the tokens a
full-resolution frame costs; 'low' is cheaper still but starts losing small
text, which is most of what screen captures are for."""
v = str(load_api_keys().get("media_resolution", "medium")).strip().lower()
return v if v in MEDIA_RESOLUTIONS else "medium"
def save_media_resolution(value: str) -> None:
v = str(value or "").strip().lower()
_save_flag("media_resolution", v if v in MEDIA_RESOLUTIONS else "medium")
def _save_flag(key: str, value) -> None:
"""Read-modify-write one key without disturbing the rest of the config."""
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
data[key] = bool(value) if isinstance(value, bool) else value
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
def get_brief_enabled() -> bool:
return load_api_keys().get("morning_brief_enabled", True)
def save_brief_enabled(enabled: bool) -> None:
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
data["morning_brief_enabled"] = enabled
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
# ── Audio devices ────────────────────────────────────────────────────────────
# Stored as device NAMES, not sounddevice indices. Indices shift every time a
# USB device is plugged in or removed, so a saved index silently starts pointing
# at a different microphone. The empty string means "system default", which is
# both the factory setting and what an unresolvable saved device falls back to —
# so unplugging a headset degrades to the built-in speakers instead of crashing.
def _patch_config(**fields) -> None:
"""Read-modify-write one or more keys in api_keys.json.
Every setter in this file open-coded this. Collapsing it here means a new
setting is one line, and there is one place where a corrupt config file is
handled instead of nine."""
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
data.update(fields)
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
def get_input_device() -> str:
"""Microphone device name, or '' for the system default."""
return (load_api_keys().get("input_device", "") or "").strip()
def save_input_device(name: str) -> None:
_patch_config(input_device=(name or "").strip())
def get_output_device() -> str:
"""Speaker device name, or '' for the system default."""
return (load_api_keys().get("output_device", "") or "").strip()
def save_output_device(name: str) -> None:
_patch_config(output_device=(name or "").strip())
def get_plugin_enabled(plugin_name: str) -> bool:
"""Plugins are enabled by default the moment they're discovered (opt-out model)."""
return load_api_keys().get("plugins_enabled", {}).get(plugin_name, True)
# ── Per-plugin settings ("tokens" / connection details) ───────────────────────
# Generic store so a plugin can declare its own config fields (PLUGIN_SETTINGS)
# and the settings UI renders + persists them WITHOUT any core edit — keeping the
# drop-in model intact. Values live under plugin_config[<namespace>][<key>].
# A namespace defaults to the plugin name, but a suite of plugins (e.g. the
# several printer plugins) can share ONE namespace.
def get_plugin_config(namespace: str) -> dict:
"""All stored values for a namespace (empty dict if none set yet)."""
cfg = load_api_keys().get("plugin_config")
val = cfg.get(namespace) if isinstance(cfg, dict) else None
return dict(val) if isinstance(val, dict) else {}
def get_plugin_setting(namespace: str, key: str, default=None):
"""A single value from a namespace, or `default` if unset."""
return get_plugin_config(namespace).get(key, default)
def save_plugin_config(namespace: str, values: dict) -> None:
"""Merge `values` into a namespace's stored config (read-modify-write, like
every other helper here). Only the provided keys are touched."""
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
pc = data.get("plugin_config")
if not isinstance(pc, dict):
pc = {}
cur = pc.get(namespace)
if not isinstance(cur, dict):
cur = {}
cur.update(values)
pc[namespace] = cur
data["plugin_config"] = pc
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
def save_plugin_enabled(plugin_name: str, enabled: bool) -> None:
ensure_config_dir()
data: dict = {}
if CONFIG_FILE.exists():
try:
data = json.loads(CONFIG_FILE.read_text(encoding="utf-8"))
except Exception:
data = {}
plugins_cfg = data.get("plugins_enabled")
if not isinstance(plugins_cfg, dict):
plugins_cfg = {}
plugins_cfg[plugin_name] = enabled
data["plugins_enabled"] = plugins_cfg
CONFIG_FILE.write_text(json.dumps(data, indent=4), encoding="utf-8")
+479
View File
@@ -0,0 +1,479 @@
import json
import re
from datetime import datetime
from threading import Lock
from pathlib import Path
import sys
def get_base_dir() -> Path:
if getattr(sys, "frozen", False):
return Path(sys.executable).parent
return Path(__file__).resolve().parent.parent
BASE_DIR = get_base_dir()
MEMORY_PATH = BASE_DIR / "memory" / "long_term.json"
_lock = Lock()
MAX_VALUE_LENGTH = 380
# ── Why there are two very different numbers here ────────────────────────────
#
# There used to be one: MEMORY_MAX_CHARS = 2200, applied to the whole store. It
# was a *storage* limit, and it existed only because the entire memory was
# pasted into the system prompt on every connect — so growing the memory grew
# every single request. When it filled, _trim_to_limit() deleted the oldest
# entries and printed one line to a console nobody reads. A memory described as
# "deeply remembers projects, preferences and personal context" was in practice
# two pages long, and quietly forgot your sister's name after a few weeks.
#
# Storage and prompt budget are now separate concerns:
#
# MEMORY_MAX_CHARS — a runaway guard, not a feature limit. Nothing normal
# reaches it; a bug writing in a loop does.
# PROMPT_CORE_CHARS — what actually rides in the system prompt every session.
# Smaller than the old whole-memory dump, so sessions
# start *faster* than before, not slower.
#
# Everything above the core stays on disk and is fetched on demand by the
# recall_memory tool — see search_memory() and format_memory_for_prompt().
MEMORY_MAX_CHARS = 200_000
PROMPT_CORE_CHARS = 900
PROMPT_INDEX_CHARS = 420
# Most entries any one category may contribute to the core block, so a person
# with forty stored preferences still gets their sister into the prompt.
PROMPT_MAX_PER_CATEGORY = 6
def _empty_memory() -> dict:
return {
"identity": {},
"preferences": {},
"projects": {},
"relationships": {},
"wishes": {},
"notes": {},
}
def load_memory() -> dict:
if not MEMORY_PATH.exists():
return _empty_memory()
with _lock:
try:
data = json.loads(MEMORY_PATH.read_text(encoding="utf-8"))
if isinstance(data, dict):
base = _empty_memory()
for key in base:
if key not in data:
data[key] = {}
return data
return _empty_memory()
except Exception as e:
print(f"[Memory] ⚠️ Load error: {e}")
return _empty_memory()
def _all_entries(memory: dict) -> list[tuple]:
entries = []
for cat, items in memory.items():
if not isinstance(items, dict):
continue
for key, entry in items.items():
if isinstance(entry, dict) and "value" in entry:
entries.append((cat, key, entry))
return entries
# Set by main.py so a trim can reach the activity log. Deleting something a
# person told you and mentioning it only on stdout is how a memory loses trust.
_trim_notifier = None
def set_trim_notifier(fn) -> None:
"""Register a callable(str) that surfaces trims to the user."""
global _trim_notifier
_trim_notifier = fn
def _trim_to_limit(memory: dict) -> dict:
if len(json.dumps(memory, ensure_ascii=False)) <= MEMORY_MAX_CHARS:
return memory
entries = _all_entries(memory)
entries.sort(key=lambda t: t[2].get("updated", "0000-00-00"))
dropped = []
for cat, key, _ in entries:
if len(json.dumps(memory, ensure_ascii=False)) <= MEMORY_MAX_CHARS:
break
del memory[cat][key]
dropped.append(f"{cat}/{key}")
print(f"[Memory] 🗑️ Trimmed {cat}/{key}")
if dropped and _trim_notifier:
try:
_trim_notifier(
f"SYS: Memory full — forgot {len(dropped)} oldest entries "
f"({', '.join(dropped[:3])}{'…' if len(dropped) > 3 else ''})"
)
except Exception:
pass
return memory
def save_memory(memory: dict) -> None:
if not isinstance(memory, dict):
return
memory = _trim_to_limit(memory)
MEMORY_PATH.parent.mkdir(parents=True, exist_ok=True)
with _lock:
MEMORY_PATH.write_text(
json.dumps(memory, indent=2, ensure_ascii=False),
encoding="utf-8",
)
def _truncate_value(val: str) -> str:
if isinstance(val, str) and len(val) > MAX_VALUE_LENGTH:
return val[:MAX_VALUE_LENGTH].rstrip() + "…"
return val
def _recursive_update(target: dict, updates: dict) -> bool:
changed = False
for key, value in updates.items():
if value is None:
continue
if isinstance(value, str) and not value.strip():
continue
if isinstance(value, dict) and "value" not in value:
if key not in target or not isinstance(target[key], dict):
target[key] = {}
changed = True
if _recursive_update(target[key], value):
changed = True
else:
new_val = _truncate_value(str(value["value"] if isinstance(value, dict) else value))
entry = {"value": new_val, "updated": datetime.now().strftime("%Y-%m-%d")}
existing = target.get(key, {})
if not isinstance(existing, dict) or existing.get("value") != new_val:
target[key] = entry
changed = True
return changed
def update_memory(memory_update: dict) -> dict:
if not isinstance(memory_update, dict) or not memory_update:
return load_memory()
memory = load_memory()
if _recursive_update(memory, memory_update):
save_memory(memory)
print(f"[Memory] 💾 Saved: {list(memory_update.keys())}")
return memory
def _entry_value(entry) -> str:
"""Accept both the {'value': ..., 'updated': ...} shape and a bare string,
because early versions of the store wrote plain strings."""
if isinstance(entry, dict):
return str(entry.get("value", "") or "").strip()
return str(entry or "").strip()
def _pretty(key: str) -> str:
return key.replace("_", " ").strip()
# Identity is always in the prompt; these categories compete for the remaining
# budget by recency.
_CATEGORY_LABELS = {
"preferences": "Preferences",
"projects": "Active projects / goals",
"relationships": "People in their life",
"wishes": "Wishes / plans",
"notes": "Notes",
}
_IDENTITY_FIELDS = ["name", "age", "birthday", "city", "job",
"language", "school", "nationality"]
def format_memory_for_prompt(memory: dict | None) -> str:
"""Build the memory block that goes into the system prompt.
This used to dump everything. It now sends three things:
1. IDENTITY - always, in full. It is small, and it is wrong for the
assistant to have to look up your name.
2. RECENT - the most recently updated entries from every other
category, up to PROMPT_CORE_CHARS. Recency is the cheapest useful
relevance signal available without embeddings.
3. AN INDEX - the *keys* of everything else, values omitted.
Point 3 is what makes recall work at all. A model cannot decide to look
something up if it does not know the thing exists: with only points 1 and 2,
"who is Ayse?" would get "I don't know" while ayse_sister sat on disk
unread. The index costs a few hundred characters and turns recall from a
gamble into a lookup.
Net effect on latency: this block is SMALLER than the old full dump, so
every session connects with fewer tokens. Occasionally the model spends one
extra round trip on recall_memory - covered by the acknowledgment it
already speaks before any slow step."""
if not memory:
return ""
core_lines: list[str] = []
# 1. Identity - always, in full
identity = memory.get("identity", {}) or {}
for field in _IDENTITY_FIELDS:
val = _entry_value(identity.get(field))
if not val:
continue
if field == "language":
# Labelled as an observation, not a setting. A bare "Language:
# English" line written months ago reads like a standing order and
# was one of the reasons a Turkish question came back in English.
core_lines.append(
f"Has spoken to you in: {val} (an observation about the past — "
f"always answer in the language of their CURRENT message)")
else:
core_lines.append(f"{field.title()}: {val}")
for key, entry in identity.items():
if key in _IDENTITY_FIELDS:
continue
val = _entry_value(entry)
if val:
core_lines.append(f"{_pretty(key).title()}: {val}")
# 2. Everything else, most recently updated first
rest: list[tuple[str, str, str, str]] = [] # (updated, cat, key, value)
for cat in _CATEGORY_LABELS:
for key, entry in (memory.get(cat, {}) or {}).items():
val = _entry_value(entry)
if not val:
continue
updated = (entry.get("updated", "") if isinstance(entry, dict) else "") or "0000-00-00"
rest.append((updated, cat, key, val))
rest.sort(key=lambda t: t[0], reverse=True)
used = sum(len(l) + 1 for l in core_lines)
shown: dict[str, list[str]] = {}
overflow: dict[str, list[str]] = {}
# Recency decides order, but no single category may take the whole budget.
# Without the cap, someone with forty stored preferences gets a prompt that
# is forty preferences and not one person's name — the categories that
# matter most in conversation are also the ones that change least often, so
# pure recency systematically buries them.
per_cat_used: dict[str, int] = {}
for _updated, cat, key, val in rest:
line = f" - {_pretty(key).title()}: {val}"
if (per_cat_used.get(cat, 0) < PROMPT_MAX_PER_CATEGORY
and used + len(line) + 1 <= PROMPT_CORE_CHARS):
shown.setdefault(cat, []).append(line)
per_cat_used[cat] = per_cat_used.get(cat, 0) + 1
used += len(line) + 1
else:
overflow.setdefault(cat, []).append(_pretty(key))
# The index is a table of contents, so it is interleaved across categories
# rather than continuing in recency order. Sorted by recency it would list
# twenty-four preferences before the first relationship, and the one entry
# the index exists for — the old fact the model has no other way to know
# about — would fall off the end.
indexed: list[str] = []
if overflow:
cats = [c for c in _CATEGORY_LABELS if overflow.get(c)]
cursor = {c: 0 for c in cats}
while cats:
for cat in list(cats):
i = cursor[cat]
if i >= len(overflow[cat]):
cats.remove(cat)
continue
indexed.append(overflow[cat][i])
cursor[cat] = i + 1
for cat, label in _CATEGORY_LABELS.items():
if shown.get(cat):
core_lines.append("")
core_lines.append(f"{label}:")
core_lines.extend(shown[cat])
if not core_lines and not indexed:
return ""
out = [
"[WHAT YOU KNOW ABOUT THIS PERSON — use naturally, never recite like a list]",
*core_lines,
]
# 3. The index of what is on disk but not in this prompt
if indexed:
budget, names = PROMPT_INDEX_CHARS, []
for n in indexed:
if budget - len(n) - 2 < 0:
break
names.append(n)
budget -= len(n) + 2
if names:
out.append("")
out.append(
"[ALSO REMEMBERED — values not shown here. Call recall_memory "
"with a keyword to read any of these before saying you do not know]"
)
out.append(", ".join(names)
+ (f" (+{len(indexed) - len(names)} more)"
if len(indexed) > len(names) else ""))
return "\n".join(out) + "\n"
# ── Recall ────────────────────────────────────────────────────────────────────
def _score(query_words: list[str], cat: str, key: str, value: str) -> int:
"""Cheap lexical relevance. No embeddings, no network, no model call - this
runs in well under a millisecond, which is the entire point: recall must
cost one model round trip, never two."""
hay_key = _pretty(key).lower()
hay_val = value.lower()
score = 0
for w in query_words:
if not w:
continue
if w == hay_key:
score += 10
elif w in hay_key:
score += 6
if w in hay_val:
score += 3
if w in cat:
score += 1
return score
def search_memory(query: str, limit: int = 8) -> str:
"""Find stored facts matching `query`. Backs the recall_memory tool.
An empty query is treated as "show me everything you know", capped - the
model asks that when the user says "what do you remember about me?"."""
memory = load_memory()
words = [w for w in re.split(r"[^\w]+", (query or "").lower()) if len(w) > 1]
rows: list[tuple[int, str, str, str]] = []
for cat, items in memory.items():
if not isinstance(items, dict):
continue # skip 'sessions', which is a list
for key, entry in items.items():
val = _entry_value(entry)
if not val:
continue
s = _score(words, cat, key, val) if words else 1
if s > 0:
rows.append((s, cat, key, val))
if not rows:
return (f"Nothing stored about '{query}'." if query
else "I have not stored anything about this person yet.")
rows.sort(key=lambda r: (-r[0], r[2]))
lines = [f"{cat}/{_pretty(key)}: {val}" for _s, cat, key, val in rows[:max(1, limit)]]
head = (f"Stored facts matching '{query}':" if query
else "Everything currently stored:")
more = (f"\n(+{len(rows) - len(lines)} more — search with a narrower keyword)"
if len(rows) > len(lines) else "")
return head + "\n" + "\n".join(lines) + more
def all_entries_for_ui() -> list[dict]:
"""Flat list for the memory panel: what JARVIS knows, and when it learned it.
Sorted newest first so the panel opens on what changed most recently."""
memory = load_memory()
rows = []
for cat, items in memory.items():
if not isinstance(items, dict):
continue
for key, entry in items.items():
val = _entry_value(entry)
if not val:
continue
rows.append({
"category": cat,
"key": key,
"value": val,
"updated": (entry.get("updated", "") if isinstance(entry, dict) else ""),
})
rows.sort(key=lambda r: (r["updated"] or "0000-00-00"), reverse=True)
return rows
def remember(key: str, value: str, category: str = "notes") -> str:
valid = {"identity", "preferences", "projects", "relationships", "wishes", "notes"}
if category not in valid:
category = "notes"
update_memory({category: {key: {"value": value}}})
return f"Remembered: {category}/{key} = {value}"
def forget(key: str, category: str = "notes") -> str:
memory = load_memory()
cat = memory.get(category, {})
if key in cat:
del cat[key]
memory[category] = cat
save_memory(memory)
return f"Forgotten: {category}/{key}"
return f"Not found: {category}/{key}"
forget_memory = forget
# ── Session memory ─────────────────────────────────────────────────────────────
_SESSION_MAX = 3 # safety cap — in practice 0-1 entries after pop
def save_session_summary(summary: str, language: str = "") -> None:
"""Append a 1-2 sentence session summary to long_term.json['sessions']."""
summary = (summary or "").strip()
if not summary:
return
memory = load_memory()
sessions = memory.get("sessions", [])
if not isinstance(sessions, list):
sessions = []
entry: dict = {
"date": datetime.now().strftime("%Y-%m-%d"),
"summary": summary[:280],
}
if language:
entry["language"] = language
sessions.append(entry)
memory["sessions"] = sessions[-_SESSION_MAX:]
with _lock:
MEMORY_PATH.parent.mkdir(parents=True, exist_ok=True)
MEMORY_PATH.write_text(
json.dumps(memory, indent=2, ensure_ascii=False),
encoding="utf-8",
)
print(f"[Memory] 📝 Session saved ({entry['date']}): {summary[:60]}…")
def pop_last_session() -> dict | None:
"""
Return AND remove the most recent session entry.
Calling this consumes the entry so it is never repeated in future briefings.
"""
with _lock:
if not MEMORY_PATH.exists():
return None
try:
memory = json.loads(MEMORY_PATH.read_text(encoding="utf-8"))
sessions = memory.get("sessions", [])
if not isinstance(sessions, list) or not sessions:
return None
entry = sessions.pop() # remove the last entry
memory["sessions"] = sessions
MEMORY_PATH.write_text(
json.dumps(memory, indent=2, ensure_ascii=False),
encoding="utf-8",
)
return entry
except Exception as e:
print(f"[Memory] ⚠️ pop_last_session error: {e}")
return None
View File
Whitespace-only changes.
+55
View File
@@ -0,0 +1,55 @@
"""Apre e chiude applicazioni, apre siti e file."""
from __future__ import annotations
import subprocess
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import osa # noqa: E402
def apri(params: dict, ctx: dict) -> str:
nome = str(params.get("nome", "")).strip()
url = str(params.get("url", "")).strip()
if url:
if not url.startswith(("http://", "https://")):
url = "https://" + url
cmd = ["open", url] if not nome else ["open", "-a", nome, url]
subprocess.run(cmd, check=False, timeout=15)
return f"Apro {url}" + (f" con {nome}." if nome else ".")
if not nome:
return "Errore: serve il nome dell'app o un indirizzo."
p = Path(nome).expanduser()
if p.exists():
subprocess.run(["open", str(p)], check=False, timeout=15)
return f"Apro {p.name}."
res = subprocess.run(["open", "-a", nome], capture_output=True, text=True, timeout=15)
if res.returncode != 0:
return f"Non trovo un'app chiamata '{nome}'."
return f"Apro {nome}."
def chiudi(params: dict, ctx: dict) -> str:
nome = str(params.get("nome", "")).strip()
if not nome:
return "Errore: serve il nome dell'app."
try:
osa('on run argv\ntell application (item 1 of argv) to quit\nend run', nome, timeout=20)
return f"Chiudo {nome}."
except Exception as err:
return f"Non riesco a chiudere {nome}: {err}"
def in_esecuzione(params: dict, ctx: dict) -> str:
out = osa('tell application "System Events" to get name of every process whose background only is false')
return "App aperte: " + out
TOOLS = [
{"name": "app_apri", "description": "Apre un'applicazione (per nome, es. Safari, Xcode), un sito (url) o un file. Se dai sia app che url, apre l'url con quell'app.",
"parameters": {"type": "object", "properties": {"nome": {"type": "string"}, "url": {"type": "string"}}}, "run": apri},
{"name": "app_chiudi", "description": "Chiude un'applicazione per nome.",
"parameters": {"type": "object", "properties": {"nome": {"type": "string"}}, "required": ["nome"]}, "run": chiudi},
{"name": "app_aperte", "description": "Elenca le applicazioni aperte.", "parameters": {"type": "object", "properties": {}}, "run": in_esecuzione},
]
+41
View File
@@ -0,0 +1,41 @@
"""Avvisi del monitor di Mail, WhatsApp e Telegram: elenco, chiusura, controllo immediato."""
from __future__ import annotations
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
def _mon():
from avatar import monitor as m
if m.monitor is None:
raise RuntimeError("Il monitor non è attivo in questo processo.")
return m.monitor
def elenca(params: dict, ctx: dict) -> str:
al = _mon().pending()
if not al:
return "Nessun avviso aperto."
return "Avvisi aperti:\n" + "\n".join(f"- [{a['id']}] {a['quando']} {a['fonte']} · {a['chi']} — {a['motivo']} ({a['priorita']})\n «{a['testo'][:160]}»" for a in al)
def chiudi(params: dict, ctx: dict) -> str:
n = _mon().close(str(params.get("id", "")).strip())
return f"Chiusi {n} avvisi." if n else "Nessun avviso corrispondente."
def controlla(params: dict, ctx: dict) -> str:
al = _mon().scan()
return f"Controllo fatto: {len(al)} nuovi avvisi." + ("" if not al else "\n" + "\n".join(f"- {a['fonte']} · {a['chi']} — {a['motivo']}" for a in al))
TOOLS = [
{"name": "attenzione_elenca", "description": "Elenca gli avvisi aperti del monitor (messaggi Mail, WhatsApp e Telegram che richiedono attenzione).",
"parameters": {"type": "object", "properties": {}}, "run": elenca},
{"name": "attenzione_chiudi", "description": "Chiude un avviso per id o per nome del mittente; senza parametri chiude tutti.",
"parameters": {"type": "object", "properties": {"id": {"type": "string"}}}, "run": chiudi},
{"name": "attenzione_controlla", "description": "Esegue subito un controllo di Mail, WhatsApp e Telegram alla ricerca di messaggi che richiedono attenzione.",
"parameters": {"type": "object", "properties": {}}, "run": controlla},
]
+78
View File
@@ -0,0 +1,78 @@
"""Browser e web: legge una pagina, la pagina aperta in Safari, cerca su YouTube o Google."""
from __future__ import annotations
import html
import re
import subprocess
import sys
from pathlib import Path
from urllib.parse import quote_plus
import requests
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import osa # noqa: E402
MAX_CHARS = 6000
UA = {"User-Agent": "Mozilla/5.0 (Macintosh) AvatarPy/1.0"}
def _html_to_text(page: str) -> str:
page = re.sub(r"<(script|style|noscript|svg|header|footer|nav)[^>]*>.*?</\1>", " ", page, flags=re.S | re.I)
page = re.sub(r"<br\s*/?>|</p>|</div>|</h\d>|</li>|</tr>", "\n", page, flags=re.I)
text = html.unescape(re.sub(r"<[^>]+>", " ", page))
lines = [" ".join(l.split()) for l in text.splitlines()]
return "\n".join(l for l in lines if len(l) > 2)
def leggi_pagina(params: dict, ctx: dict) -> str:
url = str(params.get("url", "")).strip()
if not url:
try:
url = osa('tell application "Safari" to return URL of front document')
except Exception:
return "Nessun indirizzo dato e Safari non ha pagine aperte."
if not url.startswith("http"):
url = "https://" + url
r = requests.get(url, headers=UA, timeout=15)
r.raise_for_status()
title = re.search(r"<title[^>]*>(.*?)</title>", r.text, flags=re.S | re.I)
text = _html_to_text(r.text)
return f"{html.unescape(title.group(1).strip()) if title else url}\n{text[:MAX_CHARS]}{' […]' if len(text) > MAX_CHARS else ''}"
def cerca(params: dict, ctx: dict) -> str:
q = str(params.get("query", "")).strip()
dove = str(params.get("dove", "google")).lower()
if not q:
return "Errore: serve la query."
url = {"youtube": f"https://www.youtube.com/results?search_query={quote_plus(q)}",
"maps": f"https://www.google.com/maps/search/{quote_plus(q)}",
"amazon": f"https://www.amazon.it/s?k={quote_plus(q)}",
"wikipedia": f"https://it.wikipedia.org/w/index.php?search={quote_plus(q)}"}.get(dove, f"https://www.google.com/search?q={quote_plus(q)}")
subprocess.run(["open", url], check=False, timeout=10)
return f"Apro la ricerca di «{q}» su {dove}."
def youtube(params: dict, ctx: dict) -> str:
q = str(params.get("query", "")).strip()
if not q:
return "Errore: serve cosa cercare."
r = requests.get(f"https://www.youtube.com/results?search_query={quote_plus(q)}", headers=UA, timeout=15)
m = re.search(r'"videoId":"([\w-]{11})"', r.text)
if not m:
subprocess.run(["open", f"https://www.youtube.com/results?search_query={quote_plus(q)}"], check=False)
return f"Apro i risultati YouTube per «{q}»."
subprocess.run(["open", f"https://www.youtube.com/watch?v={m.group(1)}"], check=False)
t = re.search(r'"videoId":"' + m.group(1) + r'".{0,400}?"title":\{"runs":\[\{"text":"([^"]+)"', r.text)
return f"Riproduco su YouTube: {html.unescape(t.group(1)) if t else q}."
TOOLS = [
{"name": "web_leggi_pagina", "description": "Legge il testo di una pagina web (dato l'url) o della pagina aperta in Safari, per riassumerla o rispondere.",
"parameters": {"type": "object", "properties": {"url": {"type": "string"}}}, "run": leggi_pagina},
{"name": "web_cerca_apri", "description": "Apre nel browser una ricerca su Google, YouTube, Maps, Amazon o Wikipedia.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}, "dove": {"type": "string", "enum": ["google", "youtube", "maps", "amazon", "wikipedia"]}}, "required": ["query"]}, "run": cerca},
{"name": "youtube_riproduci", "description": "Cerca un video su YouTube e apre direttamente il primo risultato.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]}, "run": youtube},
]
+33
View File
@@ -0,0 +1,33 @@
"""Riepilogo della giornata: calendario, promemoria, email non lette e meteo in un colpo solo."""
from __future__ import annotations
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
def riepilogo(params: dict, ctx: dict) -> str:
from avatar.plugins import registry
citta = str(params.get("citta", "")).strip() or "Roma"
parts = []
for label, name, args in (
("Calendario", "calendario_eventi", {"da": "oggi", "a": "domani"}),
("Promemoria", "promemoria_elenca", {"quando": "oggi"}),
("Email", "mail_non_lette", {"limite": 5}),
("Meteo", "meteo", {"citta": citta, "giorni": 1}),
):
if not registry.has(name):
continue
try:
parts.append(f"## {label}\n{registry.run(name, args)}")
except Exception as err:
parts.append(f"## {label}\nnon disponibile ({err})")
return ("\n\n".join(parts) + "\n\nFai un riepilogo parlato breve e naturale: prima gli impegni, poi i promemoria, "
"poi le email importanti (solo mittente e oggetto), infine il meteo in una frase.")
TOOLS = [
{"name": "buongiorno", "description": "Riepilogo della giornata: impegni di oggi e domani, promemoria, email non lette e meteo. Usalo per 'buongiorno', 'come si presenta la giornata', 'riepilogo'.",
"parameters": {"type": "object", "properties": {"citta": {"type": "string", "description": "Città per il meteo (default Roma)."}}}, "run": riepilogo},
]
+227
View File
@@ -0,0 +1,227 @@
"""Calendario di macOS: legge, crea, sposta e cancella eventi (EventKit)."""
from __future__ import annotations
import threading
from datetime import date, datetime, timedelta
_store = None
_lock = threading.Lock()
GIORNI = ["lun", "mar", "mer", "gio", "ven", "sab", "dom"]
MESI = ["gen", "feb", "mar", "apr", "mag", "giu", "lug", "ago", "set", "ott", "nov", "dic"]
def _get_store():
"""EKEventStore con accesso richiesto una sola volta."""
global _store
with _lock:
if _store is not None:
return _store
import EventKit
store = EventKit.EKEventStore.alloc().init()
done, res = threading.Event(), {}
def cb(granted, err):
res["granted"] = bool(granted)
done.set()
if hasattr(store, "requestFullAccessToEventsWithCompletion_"):
store.requestFullAccessToEventsWithCompletion_(cb)
else:
store.requestAccessToEntityType_completion_(EventKit.EKEntityTypeEvent, cb)
done.wait(60)
if not res.get("granted"):
raise RuntimeError("Accesso al Calendario negato: concedilo in Impostazioni di Sistema > Privacy e sicurezza > Calendari.")
_store = store
return store
def _nsdate(dt: datetime):
from Foundation import NSDate
return NSDate.dateWithTimeIntervalSince1970_(dt.timestamp())
def _pydate(nsdate) -> datetime:
return datetime.fromtimestamp(nsdate.timeIntervalSince1970())
def _parse_day(s: str) -> date:
s = (s or "").strip().lower()
today = date.today()
if s in ("", "oggi", "today"):
return today
if s in ("domani", "tomorrow"):
return today + timedelta(days=1)
if s in ("dopodomani",):
return today + timedelta(days=2)
if s in ("ieri",):
return today - timedelta(days=1)
return datetime.strptime(s[:10], "%Y-%m-%d").date()
def _parse_dt(s: str) -> datetime:
s = (s or "").strip().replace("T", " ")
for fmt in ("%Y-%m-%d %H:%M", "%Y-%m-%d %H:%M:%S", "%Y-%m-%d"):
try:
return datetime.strptime(s, fmt)
except ValueError:
continue
raise ValueError(f"Data e ora non valide: '{s}'. Usa il formato AAAA-MM-GG HH:MM.")
def _fmt(e) -> str:
start, end = _pydate(e.startDate()), _pydate(e.endDate())
day = f"{GIORNI[start.weekday()]} {start.day} {MESI[start.month - 1]}"
when = "tutto il giorno" if e.isAllDay() else f"{start:%H:%M}–{end:%H:%M}"
extra = f" @ {e.location()}" if e.location() else ""
return f"{day} {when}: {e.title()} ({e.calendar().title()}){extra}"
def _events_between(start: datetime, end: datetime):
store = _get_store()
pred = store.predicateForEventsWithStartDate_endDate_calendars_(_nsdate(start), _nsdate(end), None)
evs = list(store.eventsMatchingPredicate_(pred) or [])
evs.sort(key=lambda e: e.startDate().timeIntervalSince1970())
return evs
def eventi(params: dict, ctx: dict) -> str:
da = _parse_day(params.get("da", ""))
a = _parse_day(params.get("a", "")) if params.get("a") else da
if a < da:
a = da
evs = _events_between(datetime.combine(da, datetime.min.time()), datetime.combine(a + timedelta(days=1), datetime.min.time()))
label = f"{da:%d/%m}" if da == a else f"dal {da:%d/%m} al {a:%d/%m}"
if not evs:
return f"Nessun evento {label}."
return f"Eventi {label}:\n" + "\n".join(f"- {_fmt(e)}" for e in evs[:40])
def _find_calendar(store, name: str):
import EventKit
cals = [c for c in store.calendarsForEntityType_(EventKit.EKEntityTypeEvent) if c.allowsContentModifications()]
if name:
for c in cals:
if c.title().lower() == name.lower():
return c
for c in cals:
if name.lower() in c.title().lower():
return c
return store.defaultCalendarForNewEvents() or (cals[0] if cals else None)
def crea(params: dict, ctx: dict) -> str:
import EventKit
store = _get_store()
titolo = str(params.get("titolo", "")).strip()
if not titolo:
return "Errore: serve il titolo dell'evento."
start = _parse_dt(str(params.get("inizio", "")))
all_day = bool(params.get("tutto_il_giorno", False))
if all_day:
start = start.replace(hour=0, minute=0)
end = start + timedelta(days=1)
elif params.get("fine"):
end = _parse_dt(str(params["fine"]))
else:
end = start + timedelta(minutes=int(params.get("durata_minuti") or 60))
ev = EventKit.EKEvent.eventWithEventStore_(store)
ev.setTitle_(titolo)
ev.setStartDate_(_nsdate(start))
ev.setEndDate_(_nsdate(end))
ev.setAllDay_(all_day)
cal = _find_calendar(store, str(params.get("calendario", "")))
if cal is None:
return "Errore: nessun calendario modificabile disponibile."
ev.setCalendar_(cal)
if params.get("luogo"):
ev.setLocation_(str(params["luogo"]))
if params.get("note"):
ev.setNotes_(str(params["note"]))
ok, err = store.saveEvent_span_error_(ev, EventKit.EKSpanThisEvent, None)
if not ok:
return f"Errore nel salvataggio: {err}"
return f"Evento creato: {_fmt(ev)}"
def _match(params: dict) -> tuple[list, str]:
titolo = str(params.get("titolo", "")).strip().lower()
giorno = _parse_day(params.get("giorno", ""))
evs = _events_between(datetime.combine(giorno, datetime.min.time()), datetime.combine(giorno + timedelta(days=1), datetime.min.time()))
found = [e for e in evs if titolo and titolo in (e.title() or "").lower()] if titolo else evs
return found, f"{giorno:%d/%m}"
def elimina(params: dict, ctx: dict) -> str:
import EventKit
found, label = _match(params)
if not found:
return f"Nessun evento corrispondente il {label}."
if len(found) > 1:
return "Ho trovato più eventi, dimmi quale:\n" + "\n".join(f"- {_fmt(e)}" for e in found)
ev = found[0]
desc = _fmt(ev)
def do() -> str:
store = _get_store()
ok, err = store.removeEvent_span_error_(ev, EventKit.EKSpanThisEvent, None)
msg = f"Evento cancellato: {desc}" if ok else f"Errore nella cancellazione: {err}"
if ctx.get("say"):
ctx["say"](msg)
return msg
confirm = ctx.get("confirm")
if confirm:
return confirm("calendario_elimina", "Cancellare l'evento?", desc, do)
# Senza interfaccia (server MCP): serve la conferma esplicita nei parametri.
if str(params.get("confermato", "")).lower() in ("true", "1", "sì", "si", "yes"):
return do()
return f"Prima di cancellare chiedi conferma all'utente per: {desc}. Poi richiama con confermato=true."
def sposta(params: dict, ctx: dict) -> str:
import EventKit
found, label = _match(params)
if not found:
return f"Nessun evento corrispondente il {label}."
if len(found) > 1:
return "Ho trovato più eventi, dimmi quale:\n" + "\n".join(f"- {_fmt(e)}" for e in found)
ev = found[0]
new_start = _parse_dt(str(params.get("nuovo_inizio", "")))
duration = _pydate(ev.endDate()) - _pydate(ev.startDate())
ev.setStartDate_(_nsdate(new_start))
ev.setEndDate_(_nsdate(new_start + duration))
ok, err = _get_store().saveEvent_span_error_(ev, EventKit.EKSpanThisEvent, None)
return f"Evento spostato: {_fmt(ev)}" if ok else f"Errore: {err}"
DAY = {"type": "string", "description": "Giorno: 'oggi', 'domani', 'dopodomani' oppure AAAA-MM-GG."}
TOOLS = [
{"name": "calendario_eventi",
"description": "Elenca gli eventi del Calendario di macOS in un giorno o in un intervallo. Usalo per domande come 'che impegni ho', 'cosa ho domani', 'sono libero venerdì'.",
"parameters": {"type": "object", "properties": {"da": DAY, "a": {**DAY, "description": "Ultimo giorno dell'intervallo (facoltativo)."}}, "required": ["da"]},
"run": eventi},
{"name": "calendario_crea",
"description": "Crea un evento nel Calendario. Converti tu le espressioni come 'domani alle 15' in data e ora assolute usando la data odierna del prompt.",
"parameters": {"type": "object", "properties": {
"titolo": {"type": "string"},
"inizio": {"type": "string", "description": "Inizio, formato AAAA-MM-GG HH:MM."},
"fine": {"type": "string", "description": "Fine, formato AAAA-MM-GG HH:MM (facoltativo)."},
"durata_minuti": {"type": "integer", "description": "Durata in minuti se manca la fine (default 60)."},
"tutto_il_giorno": {"type": "boolean"},
"calendario": {"type": "string", "description": "Nome del calendario (facoltativo, altrimenti quello predefinito)."},
"luogo": {"type": "string"}, "note": {"type": "string"}},
"required": ["titolo", "inizio"]},
"run": crea},
{"name": "calendario_sposta",
"description": "Sposta un evento esistente a un nuovo orario, mantenendone la durata.",
"parameters": {"type": "object", "properties": {"titolo": {"type": "string", "description": "Parte del titolo."}, "giorno": DAY,
"nuovo_inizio": {"type": "string", "description": "Nuovo inizio, AAAA-MM-GG HH:MM."}},
"required": ["titolo", "giorno", "nuovo_inizio"]},
"run": sposta},
{"name": "calendario_elimina",
"description": "Cancella un evento. Azione irreversibile: viene chiesta conferma sullo schermo; se lo strumento risponde con [CONFIRMATION_PENDING], di' all'utente di confermare sul pannello e non dire che è fatto.",
"parameters": {"type": "object", "properties": {"titolo": {"type": "string", "description": "Parte del titolo."}, "giorno": DAY,
"confermato": {"type": "boolean", "description": "Solo senza interfaccia: true dopo la conferma esplicita dell'utente."}},
"required": ["titolo", "giorno"]},
"run": elimina},
]
+40
View File
@@ -0,0 +1,40 @@
"""Comandi Rapidi (Shortcuts) di macOS: elenca ed esegue i tuoi comandi, anche per la domotica HomeKit."""
from __future__ import annotations
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import sh # noqa: E402
def elenca(params: dict, ctx: dict) -> str:
names = [n for n in sh("shortcuts", "list").splitlines() if n.strip()]
q = str(params.get("filtro", "")).strip().lower()
if q:
names = [n for n in names if q in n.lower()]
if not names:
return "Nessun comando rapido" + (f" contenente '{q}'." if q else ".")
return f"Comandi rapidi ({len(names)}):\n" + "\n".join(f"- {n}" for n in names[:40])
def esegui(params: dict, ctx: dict) -> str:
nome = str(params.get("nome", "")).strip()
if not nome:
return "Errore: serve il nome del comando rapido."
names = [n for n in sh("shortcuts", "list").splitlines() if n.strip()]
match = next((n for n in names if n.lower() == nome.lower()), None) or next((n for n in names if nome.lower() in n.lower()), None)
if not match:
return f"Nessun comando rapido chiamato '{nome}'."
inp = str(params.get("input", "")).strip()
out = sh("shortcuts", "run", match, *(["-i", "-"] if inp else []), timeout=120) if not inp else \
__import__("subprocess").run(["shortcuts", "run", match, "-i", "-"], input=inp, capture_output=True, text=True, timeout=120).stdout
return f"Eseguito «{match}»." + (f" Risultato: {out.strip()[:500]}" if out and out.strip() else "")
TOOLS = [
{"name": "comandi_rapidi_elenca", "description": "Elenca i Comandi Rapidi disponibili sul Mac (anche quelli per luci, prese e casa).",
"parameters": {"type": "object", "properties": {"filtro": {"type": "string"}}}, "run": elenca},
{"name": "comandi_rapidi_esegui", "description": "Esegue un Comando Rapido per nome (es. 'Accendi luci soggiorno'). Usalo per domotica e automazioni create dall'utente.",
"parameters": {"type": "object", "properties": {"nome": {"type": "string"}, "input": {"type": "string", "description": "Testo da passare in ingresso (facoltativo)."}}, "required": ["nome"]}, "run": esegui},
]
+65
View File
@@ -0,0 +1,65 @@
"""Contatti di macOS: cerca numero, email e dati di una persona (framework Contacts)."""
from __future__ import annotations
import threading
_store = None
_lock = threading.Lock()
def _get_store():
global _store
with _lock:
if _store is not None:
return _store
import Contacts
store = Contacts.CNContactStore.alloc().init()
done, res = threading.Event(), {}
def cb(granted, err):
res["granted"] = bool(granted)
done.set()
store.requestAccessForEntityType_completionHandler_(Contacts.CNEntityTypeContacts, cb)
done.wait(60)
if not res.get("granted"):
raise RuntimeError("Accesso ai Contatti negato: concedilo in Impostazioni di Sistema > Privacy e sicurezza > Contatti.")
_store = store
return store
def cerca(params: dict, ctx: dict) -> str:
import Contacts
q = str(params.get("nome", "")).strip()
if not q:
return "Errore: serve il nome."
store = _get_store()
keys = [Contacts.CNContactGivenNameKey, Contacts.CNContactFamilyNameKey, Contacts.CNContactOrganizationNameKey,
Contacts.CNContactPhoneNumbersKey, Contacts.CNContactEmailAddressesKey]
pred = Contacts.CNContact.predicateForContactsMatchingName_(q)
people, err = store.unifiedContactsMatchingPredicate_keysToFetch_error_(pred, keys, None)
if err:
return f"Errore: {err}"
people = list(people or [])[:6]
if not people:
return f"Nessun contatto trovato per '{q}'."
lines = []
for p in people:
name = f"{p.givenName()} {p.familyName()}".strip() or p.organizationName()
bits = [name]
tels = [ph.value().stringValue() for ph in p.phoneNumbers()]
mails = [str(e.value()) for e in p.emailAddresses()]
if tels:
bits.append("tel " + ", ".join(tels))
if mails:
bits.append("email " + ", ".join(mails))
if p.organizationName() and p.organizationName() != name:
bits.append(p.organizationName())
lines.append("- " + " · ".join(bits))
return "\n".join(lines)
TOOLS = [
{"name": "contatti_cerca", "description": "Cerca una persona in Contatti e restituisce telefono, email e azienda. Usalo per trovare l'indirizzo email o il numero prima di scrivere a qualcuno per nome.",
"parameters": {"type": "object", "properties": {"nome": {"type": "string"}}, "required": ["nome"]}, "run": cerca},
]
+100
View File
@@ -0,0 +1,100 @@
"""File e documenti: cerca con Spotlight, elenca cartelle, legge testo, PDF e Word."""
from __future__ import annotations
import subprocess
from datetime import datetime
from pathlib import Path
HOME = Path.home()
MAX_CHARS = 6000
FOLDERS = {"desktop": HOME / "Desktop", "scrivania": HOME / "Desktop", "documenti": HOME / "Documents", "download": HOME / "Downloads",
"immagini": HOME / "Pictures", "musica": HOME / "Music", "home": HOME}
def _folder(name: str) -> Path:
n = (name or "").strip().lower()
if n in FOLDERS:
return FOLDERS[n]
p = Path(name).expanduser()
return p if p.exists() else HOME
def cerca(params: dict, ctx: dict) -> str:
q = str(params.get("query", "")).strip()
if not q:
return "Errore: serve il campo 'query'."
folder = _folder(str(params.get("cartella", "home")))
kind = str(params.get("tipo", "")).lower()
query = f'kMDItemFSName == "*{q}*"cd || kMDItemTextContent == "*{q}*"cd'
if kind == "pdf":
query = f'({query}) && kMDItemContentType == "com.adobe.pdf"'
elif kind in ("immagine", "immagini"):
query = f'({query}) && kMDItemContentTypeTree == "public.image"'
elif kind in ("documento", "documenti"):
query = f'({query}) && (kMDItemContentTypeTree == "public.text" || kMDItemContentType == "com.adobe.pdf" || kMDItemContentType == "org.openxmlformats.wordprocessingml.document")'
res = subprocess.run(["mdfind", "-onlyin", str(folder), query], capture_output=True, text=True, timeout=30)
paths = [Path(p) for p in res.stdout.splitlines() if p and "/Library/" not in p][:200]
if not paths:
return f"Nessun file trovato per '{q}' in {folder.name or 'home'}."
paths.sort(key=lambda p: p.stat().st_mtime if p.exists() else 0, reverse=True)
lines = [f"- {p.name} ({datetime.fromtimestamp(p.stat().st_mtime):%d/%m/%Y}) — {p.parent}" for p in paths[:12] if p.exists()]
return f"Trovati {len(paths)} file per '{q}':\n" + "\n".join(lines)
def elenca(params: dict, ctx: dict) -> str:
folder = _folder(str(params.get("cartella", "desktop")))
items = sorted([p for p in folder.iterdir() if not p.name.startswith(".")], key=lambda p: p.stat().st_mtime, reverse=True)
if not items:
return f"La cartella {folder.name} è vuota."
lines = [f"- {p.name}{'/' if p.is_dir() else ''} ({datetime.fromtimestamp(p.stat().st_mtime):%d/%m %H:%M})" for p in items[:25]]
return f"{folder.name} ({len(items)} elementi, i più recenti):\n" + "\n".join(lines)
def _read(path: Path) -> str:
ext = path.suffix.lower()
if ext == ".pdf":
from pypdf import PdfReader
reader = PdfReader(str(path))
return "\n".join((page.extract_text() or "") for page in reader.pages[:30])
if ext == ".docx":
import docx
d = docx.Document(str(path))
return "\n".join(p.text for p in d.paragraphs)
if ext in (".txt", ".md", ".csv", ".json", ".py", ".js", ".ts", ".html", ".rtf", ".log"):
return path.read_text(encoding="utf-8", errors="replace")
try:
return subprocess.run(["textutil", "-convert", "txt", "-stdout", str(path)], capture_output=True, text=True, timeout=30).stdout
except Exception:
return ""
def leggi(params: dict, ctx: dict) -> str:
name = str(params.get("percorso", "")).strip()
if not name:
return "Errore: serve il percorso o il nome del file."
path = Path(name).expanduser()
if not path.exists():
res = subprocess.run(["mdfind", "-onlyin", str(HOME), f'kMDItemFSName == "*{path.name}*"cd'], capture_output=True, text=True, timeout=20)
cands = [Path(p) for p in res.stdout.splitlines() if p and "/Library/" not in p]
if not cands:
return f"File non trovato: {name}"
cands.sort(key=lambda p: p.stat().st_mtime, reverse=True)
path = cands[0]
try:
text = _read(path)
except Exception as err:
return f"Non riesco a leggere {path.name}: {err}"
text = " ".join(text.split())
if not text:
return f"{path.name}: nessun testo estraibile."
return f"{path.name} ({len(text)} caratteri):\n{text[:MAX_CHARS]}{' […]' if len(text) > MAX_CHARS else ''}"
TOOLS = [
{"name": "file_cerca", "description": "Cerca file per nome o contenuto con Spotlight, in tutta la home o in una cartella (desktop, documenti, download…). Tipo: pdf, immagine, documento.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}, "cartella": {"type": "string"}, "tipo": {"type": "string"}}, "required": ["query"]}, "run": cerca},
{"name": "file_elenca", "description": "Elenca i file di una cartella (desktop, documenti, download o un percorso), i più recenti prima.",
"parameters": {"type": "object", "properties": {"cartella": {"type": "string"}}}, "run": elenca},
{"name": "file_leggi", "description": "Legge il testo di un file (txt, md, PDF, Word, pagine, rtf) per riassumerlo o rispondere a domande. Accetta un percorso o solo il nome (lo cerca).",
"parameters": {"type": "object", "properties": {"percorso": {"type": "string"}}, "required": ["percorso"]}, "run": leggi},
]
+126
View File
@@ -0,0 +1,126 @@
"""Controlli del Mac: volume, luminosità, non disturbare, Wi-Fi, batteria, disco, schermo, screenshot."""
from __future__ import annotations
import re
import subprocess
import sys
from datetime import datetime
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import osa, sh, confirm_or_param, CONFERMATO # noqa: E402
def volume(params: dict, ctx: dict) -> str:
az = str(params.get("azione", "imposta")).lower()
cur = int(osa("output volume of (get volume settings)"))
if az == "leggi":
muted = osa("output muted of (get volume settings)") == "true"
return f"Volume al {cur}%" + (" (silenziato)." if muted else ".")
if az == "muto":
osa("set volume with output muted"); return "Audio silenziato."
if az == "riattiva":
osa("set volume without output muted"); return "Audio riattivato."
v = int(params.get("valore") or 0)
if az == "alza":
v = min(100, cur + (int(params.get("valore") or 10)))
elif az == "abbassa":
v = max(0, cur - (int(params.get("valore") or 10)))
v = max(0, min(100, v))
osa(f"set volume output volume {v}")
return f"Volume al {v}%."
def luminosita(params: dict, ctx: dict) -> str:
az = str(params.get("azione", "alza")).lower()
key = 144 if az == "alza" else 145
n = max(1, min(16, int(params.get("passi") or 4)))
for _ in range(n):
osa(f'tell application "System Events" to key code {key}')
return f"Luminosità {'aumentata' if az == 'alza' else 'diminuita'}."
def non_disturbare(params: dict, ctx: dict) -> str:
on = bool(params.get("attiva", True))
try:
sh("shortcuts", "run", "Non disturbare ON" if on else "Non disturbare OFF")
return "Non disturbare " + ("attivato." if on else "disattivato.")
except Exception:
return ("Per comandare Non disturbare crea due Comandi Rapidi chiamati 'Non disturbare ON' e 'Non disturbare OFF' "
"con l'azione 'Imposta full immersion'. Poi funzionerà a voce.")
def wifi(params: dict, ctx: dict) -> str:
on = bool(params.get("attiva", True))
try:
dev = sh("networksetup", "-listallhardwareports")
m = re.search(r"Hardware Port: Wi-Fi\nDevice: (\w+)", dev)
iface = m.group(1) if m else "en0"
sh("networksetup", "-setairportpower", iface, "on" if on else "off")
return "Wi-Fi " + ("acceso." if on else "spento.")
except Exception as err:
return f"Non riesco a cambiare il Wi-Fi: {err}"
def stato(params: dict, ctx: dict) -> str:
parts = []
try:
batt = sh("pmset", "-g", "batt")
m = re.search(r"(\d+)%;\s*([^;]+);", batt)
if m:
st = {"charging": "in carica", "discharging": "in uso a batteria", "charged": "carica", "AC attached": "collegato"}.get(m.group(2).strip(), m.group(2).strip())
parts.append(f"Batteria al {m.group(1)}%, {st}")
except Exception:
pass
try:
df = sh("df", "-h", "/System/Volumes/Data").splitlines()[-1].split()
parts.append(f"Disco: {df[3]} liberi su {df[1]} ({df[4]} usato)")
except Exception:
pass
try:
up = sh("uptime")
m = re.search(r"up\s+([^,]+(?:,\s*\d+:\d+)?)", up)
if m:
parts.append(f"Acceso da {m.group(1).strip()}")
except Exception:
pass
try:
wifi_name = sh("networksetup", "-getairportnetwork", "en0")
parts.append(wifi_name.replace("Current Wi-Fi Network:", "Wi-Fi:").strip())
except Exception:
pass
return ". ".join(parts) + "." if parts else "Stato non disponibile."
def schermo(params: dict, ctx: dict) -> str:
az = str(params.get("azione", "spegni")).lower()
if az == "spegni":
sh("pmset", "displaysleepnow"); return "Schermo spento."
if az == "blocca":
osa('tell application "System Events" to keystroke "q" using {command down, control down}'); return "Mac bloccato."
if az == "sospendi":
return confirm_or_param(ctx, params, "mac_sospendi", "Mettere il Mac in stop?", "Il Mac andrà in stop.", lambda: (sh("pmset", "sleepnow"), "Stop.")[1])
return "Azione sconosciuta: spegni, blocca o sospendi."
def screenshot(params: dict, ctx: dict) -> str:
dest = Path.home() / "Desktop" / f"Screenshot {datetime.now():%Y-%m-%d %H.%M.%S}.png"
sh("screencapture", "-x", str(dest))
return f"Screenshot salvato sul Desktop: {dest.name}"
TOOLS = [
{"name": "mac_volume", "description": "Volume di sistema: leggi, imposta (0-100), alza, abbassa, muto, riattiva.",
"parameters": {"type": "object", "properties": {"azione": {"type": "string", "enum": ["leggi", "imposta", "alza", "abbassa", "muto", "riattiva"]}, "valore": {"type": "integer"}}, "required": ["azione"]}, "run": volume},
{"name": "mac_luminosita", "description": "Alza o abbassa la luminosità dello schermo (a passi).",
"parameters": {"type": "object", "properties": {"azione": {"type": "string", "enum": ["alza", "abbassa"]}, "passi": {"type": "integer"}}, "required": ["azione"]}, "run": luminosita},
{"name": "mac_non_disturbare", "description": "Attiva o disattiva Non disturbare (full immersion).",
"parameters": {"type": "object", "properties": {"attiva": {"type": "boolean"}}, "required": ["attiva"]}, "run": non_disturbare},
{"name": "mac_wifi", "description": "Accende o spegne il Wi-Fi.",
"parameters": {"type": "object", "properties": {"attiva": {"type": "boolean"}}, "required": ["attiva"]}, "run": wifi},
{"name": "mac_stato", "description": "Stato del Mac: batteria, spazio su disco, tempo di accensione, rete Wi-Fi.",
"parameters": {"type": "object", "properties": {}}, "run": stato},
{"name": "mac_schermo", "description": "Spegne lo schermo, blocca il Mac o lo mette in stop (con conferma).",
"parameters": {"type": "object", "properties": {"azione": {"type": "string", "enum": ["spegni", "blocca", "sospendi"]}, "confermato": CONFERMATO}, "required": ["azione"]}, "run": schermo},
{"name": "mac_screenshot", "description": "Cattura lo schermo e salva l'immagine sul Desktop.", "parameters": {"type": "object", "properties": {}}, "run": screenshot},
]
+223
View File
@@ -0,0 +1,223 @@
"""Mail di macOS: posta non letta, ricerca, lettura (dall'indice di Mail) e invio (AppleScript)."""
from __future__ import annotations
import html
import re
import sqlite3
import subprocess
from datetime import datetime
from pathlib import Path
MAIL_DIR = Path.home() / "Library" / "Mail"
INDEX = MAIL_DIR / "V10" / "MailData" / "Envelope Index"
MAX_BODY = 1800
SELECT = """
SELECT m.ROWID, COALESCE(a.comment, ''), COALESCE(a.address, ''), COALESCE(s.subject, ''),
m.date_received, m.read, mb.url, COALESCE(su.summary, '')
FROM messages m
JOIN mailboxes mb ON m.mailbox = mb.ROWID
LEFT JOIN addresses a ON m.sender = a.ROWID
LEFT JOIN subjects s ON m.subject = s.ROWID
LEFT JOIN summaries su ON m.summary = su.ROWID
WHERE m.deleted = 0 AND (mb.url LIKE '%/INBOX' OR mb.url LIKE '%/Inbox')
"""
def _query(where: str, params: tuple, limit: int) -> list[dict]:
if not INDEX.exists():
raise RuntimeError("Indice di Mail non trovato: l'app Mail è configurata su questo Mac?")
con = sqlite3.connect(f"file:{INDEX}?mode=ro", uri=True, timeout=5)
try:
rows = con.execute(f"{SELECT} {where} ORDER BY m.date_received DESC LIMIT ?", (*params, limit)).fetchall()
finally:
con.close()
out = []
for rid, name, addr, subj, ts, read, url, summary in rows:
who = f"{name} <{addr}>" if name and addr and name != addr else (name or addr or "?")
out.append({"id": rid, "da": who, "oggetto": subj or "(senza oggetto)", "data": datetime.fromtimestamp(ts or 0),
"letta": bool(read), "url": url, "anteprima": " ".join(str(summary).split())[:120]})
return out
def _fmt(rows: list[dict], preview: bool = False) -> str:
lines = []
for r in rows:
line = f"- [{r['id']}] {'' if r['letta'] else '● '}{r['data']:%d/%m %H:%M} — {r['da']}: {r['oggetto']}"
if preview and r["anteprima"]:
line += f"\n {r['anteprima']}"
lines.append(line)
return "\n".join(lines)
def non_lette(params: dict, ctx: dict) -> str:
lim = max(1, min(int(params.get("limite") or 10), 30))
rows = _query("AND m.read = 0", (), lim)
if not rows:
return "Nessuna email non letta nella posta in arrivo."
return f"Email non lette (le più recenti, max {lim}):\n{_fmt(rows, preview=True)}\nPer leggerne una usa mail_leggi con l'id tra parentesi."
def cerca(params: dict, ctx: dict) -> str:
q = str(params.get("query", "")).strip()
lim = max(1, min(int(params.get("limite") or 10), 30))
if not q:
return "Errore: serve il campo 'query'."
like = f"%{q}%"
rows = _query("AND (s.subject LIKE ? OR a.address LIKE ? OR a.comment LIKE ?)", (like, like, like), lim)
return f"Risultati per '{q}':\n{_fmt(rows)}\nPer leggerne una usa mail_leggi con l'id." if rows else f"Nessuna email trovata per '{q}'."
# ── Lettura dal file .emlx ────────────────────────────────────────────────────
def _mailbox_dir(url: str) -> Path | None:
m = re.match(r"^[a-z]+://([^/]+)/(.+)$", url)
if not m:
return None
account, path = m.group(1), m.group(2)
d = MAIL_DIR / "V10" / account
for part in path.split("/"):
d = d / f"{part}.mbox"
return d if d.exists() else None
def _find_emlx(rid: int, url: str) -> Path | None:
base = _mailbox_dir(url)
if base is None:
return None
digits = list(str(rid // 1000))[::-1] if rid >= 1000 else []
for sub in base.iterdir():
if not sub.is_dir():
continue
data = sub / "Data"
cand = data.joinpath(*digits) / "Messages" if digits else data / "Messages"
for name in (f"{rid}.emlx", f"{rid}.partial.emlx"):
p = cand / name
if p.exists():
return p
for p in base.rglob(f"{rid}*.emlx"):
return p
return None
def _emlx_message(path: Path):
import email
from email import policy
data = path.read_bytes()
nl = data.find(b"\n")
try:
n = int(data[:nl].strip())
raw = data[nl + 1:nl + 1 + n]
except ValueError:
raw = data
return email.message_from_bytes(raw, policy=policy.default)
def _body_text(msg) -> str:
try:
part = msg.get_body(preferencelist=("plain", "html"))
except Exception:
part = None
if part is None:
return ""
try:
text = part.get_content()
except Exception:
return ""
if part.get_content_type() == "text/html":
text = re.sub(r"<(script|style)[^>]*>.*?</\1>", " ", text, flags=re.S | re.I)
text = re.sub(r"<br\s*/?>|</p>|</div>|</tr>|</li>", "\n", text, flags=re.I)
text = html.unescape(re.sub(r"<[^>]+>", " ", text))
return " ".join(text.split())
def leggi(params: dict, ctx: dict) -> str:
mid = str(params.get("id", "")).strip()
if not mid.isdigit():
return "Errore: serve l'id numerico del messaggio preso dall'elenco."
rows = _query("AND m.ROWID = ?", (int(mid),), 1)
if not rows:
return "Messaggio non trovato."
r = rows[0]
path = _find_emlx(r["id"], r["url"])
body = ""
if path is not None:
try:
body = _body_text(_emlx_message(path))
except Exception as err:
body = f"(testo non leggibile: {err})"
if not body:
body = r["anteprima"] or "(testo non disponibile: il messaggio non è ancora scaricato in locale)"
if len(body) > MAX_BODY:
body = body[:MAX_BODY] + " […]"
return f"Da: {r['da']}\nOggetto: {r['oggetto']}\nData: {r['data']:%d/%m/%Y %H:%M}\n\n{body}"
# ── Invio via AppleScript ─────────────────────────────────────────────────────
SEND_SCRIPT = '''
on run argv
set addr to item 1 of argv
set subj to item 2 of argv
set body to item 3 of argv
tell application "Mail"
set msg to make new outgoing message with properties {subject:subj, content:body, visible:false}
tell msg to make new to recipient at end of to recipients with properties {address:addr}
send msg
end tell
return "ok"
end run
'''
def _osa(script: str, *args: str, timeout: int = 60) -> str:
res = subprocess.run(["osascript", "-e", script, "--", *args], capture_output=True, text=True, timeout=timeout)
if res.returncode != 0:
err = (res.stderr or "").strip()
if "-1743" in err:
raise RuntimeError("Accesso a Mail negato: concedilo in Impostazioni di Sistema > Privacy e sicurezza > Automazione.")
raise RuntimeError(err.splitlines()[-1] if err else "errore AppleScript")
return res.stdout.rstrip("\n")
def invia(params: dict, ctx: dict) -> str:
a = str(params.get("a", "")).strip()
oggetto = str(params.get("oggetto", "")).strip()
testo = str(params.get("testo", "")).strip()
if "@" not in a or not oggetto or not testo:
return "Errore: servono destinatario (indirizzo email), oggetto e testo."
desc = f"A: {a}\nOggetto: {oggetto}\n{testo[:160]}{'…' if len(testo) > 160 else ''}"
def do() -> str:
_osa(SEND_SCRIPT, a, oggetto, testo)
msg = f"Email inviata a {a} con oggetto «{oggetto}»."
if ctx.get("say"):
ctx["say"](msg)
return msg
confirm = ctx.get("confirm")
if confirm:
return confirm("mail_invia", "Inviare questa email?", desc, do)
if str(params.get("confermato", "")).lower() in ("true", "1", "sì", "si", "yes"):
return do()
return f"Prima di inviare chiedi conferma all'utente leggendo destinatario, oggetto e testo:\n{desc}\nPoi richiama con confermato=true."
TOOLS = [
{"name": "mail_non_lette",
"description": "Elenca le email non lette nella posta in arrivo di Mail (mittente, oggetto, data, anteprima, id). Usalo per 'ho email nuove?', 'leggimi la posta'.",
"parameters": {"type": "object", "properties": {"limite": {"type": "integer", "description": "Quante al massimo (default 10)."}}},
"run": non_lette},
{"name": "mail_cerca",
"description": "Cerca email nella posta in arrivo per parola nell'oggetto o nel mittente (nome o indirizzo).",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}, "limite": {"type": "integer"}}, "required": ["query"]},
"run": cerca},
{"name": "mail_leggi",
"description": "Legge il contenuto di un'email dato il suo id (dall'elenco). Riassumi il testo a voce, non leggerlo integralmente se è lungo.",
"parameters": {"type": "object", "properties": {"id": {"type": "string", "description": "Id numerico del messaggio."}}, "required": ["id"]},
"run": leggi},
{"name": "mail_invia",
"description": "Invia un'email dall'account predefinito di Mail. Azione irreversibile: viene chiesta conferma sullo schermo; se lo strumento risponde con [CONFIRMATION_PENDING], di' all'utente di confermare sul pannello e non dire che è inviata. Scrivi tu il testo completo in italiano, con saluto e firma con il nome dell'utente.",
"parameters": {"type": "object", "properties": {"a": {"type": "string", "description": "Indirizzo email del destinatario."}, "oggetto": {"type": "string"}, "testo": {"type": "string"},
"confermato": {"type": "boolean", "description": "Solo senza interfaccia: true dopo la conferma esplicita dell'utente."}},
"required": ["a", "oggetto", "testo"]},
"run": invia},
]
+41
View File
@@ -0,0 +1,41 @@
"""Messaggi (iMessage/SMS): invia un messaggio a un numero o a un contatto, con conferma."""
from __future__ import annotations
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import osa, confirm_or_param, CONFERMATO # noqa: E402
SEND = '''on run argv
tell application "Messages"
set svc to 1st account whose service type = iMessage
set b to participant (item 1 of argv) of svc
send (item 2 of argv) to b
end tell
return "ok"
end run'''
def invia(params: dict, ctx: dict) -> str:
a = str(params.get("a", "")).strip()
testo = str(params.get("testo", "")).strip()
if not a or not testo:
return "Errore: servono destinatario (numero con prefisso, es. +39…, o email iMessage) e testo."
if "@" not in a and not a.replace("+", "").replace(" ", "").isdigit():
return "Il destinatario deve essere un numero di telefono o un'email: cerca il contatto con contatti_cerca."
def do() -> str:
osa(SEND, a.replace(" ", ""), testo, timeout=30)
msg = f"Messaggio inviato a {a}."
if ctx.get("say"):
ctx["say"](msg)
return msg
return confirm_or_param(ctx, params, "messaggi_invia", "Inviare questo messaggio?", f"A: {a}\n{testo[:200]}", do)
TOOLS = [
{"name": "messaggi_invia", "description": "Invia un iMessage a un numero di telefono (con prefisso) o a un'email. Azione irreversibile con conferma sullo schermo; con [CONFIRMATION_PENDING] di' all'utente di confermare e non dire che è inviato. Per i nomi usa prima contatti_cerca.",
"parameters": {"type": "object", "properties": {"a": {"type": "string"}, "testo": {"type": "string"}, "confermato": CONFERMATO}, "required": ["a", "testo"]}, "run": invia},
]
+50
View File
@@ -0,0 +1,50 @@
"""Meteo con Open-Meteo (gratuito, senza chiave): oggi, domani o i prossimi giorni per una città."""
from __future__ import annotations
from datetime import datetime
import requests
CODES = {0: "sereno", 1: "prevalentemente sereno", 2: "parzialmente nuvoloso", 3: "coperto", 45: "nebbia", 48: "nebbia con brina",
51: "pioviggine leggera", 53: "pioviggine", 55: "pioviggine intensa", 61: "pioggia leggera", 63: "pioggia", 65: "pioggia forte",
71: "neve leggera", 73: "neve", 75: "neve forte", 80: "rovesci leggeri", 81: "rovesci", 82: "rovesci violenti", 95: "temporale",
96: "temporale con grandine", 99: "temporale con grandine forte"}
GIORNI = ["lunedì", "martedì", "mercoledì", "giovedì", "venerdì", "sabato", "domenica"]
def _geocode(city: str) -> tuple[float, float, str]:
r = requests.get("https://geocoding-api.open-meteo.com/v1/search", params={"name": city, "count": 1, "language": "it"}, timeout=10)
r.raise_for_status()
res = (r.json().get("results") or [])
if not res:
raise ValueError(f"Città non trovata: {city}")
g = res[0]
return g["latitude"], g["longitude"], f"{g['name']}{', ' + g['admin1'] if g.get('admin1') else ''}"
def previsioni(params: dict, ctx: dict) -> str:
city = str(params.get("citta", "")).strip() or "Roma"
giorni = max(1, min(7, int(params.get("giorni") or 2)))
lat, lon, label = _geocode(city)
r = requests.get("https://api.open-meteo.com/v1/forecast", params={
"latitude": lat, "longitude": lon, "timezone": "auto",
"current": "temperature_2m,apparent_temperature,weather_code,wind_speed_10m,relative_humidity_2m",
"daily": "weather_code,temperature_2m_max,temperature_2m_min,precipitation_probability_max,precipitation_sum",
"forecast_days": giorni}, timeout=10)
r.raise_for_status()
d = r.json()
cur = d.get("current", {})
out = [f"Meteo a {label}: ora {CODES.get(cur.get('weather_code'), '')}, {cur.get('temperature_2m')}° (percepiti {cur.get('apparent_temperature')}°), vento {cur.get('wind_speed_10m')} km/h, umidità {cur.get('relative_humidity_2m')}%."]
daily = d.get("daily", {})
for i, day in enumerate(daily.get("time", [])):
dt = datetime.strptime(day, "%Y-%m-%d")
nome = "oggi" if i == 0 else ("domani" if i == 1 else GIORNI[dt.weekday()])
out.append(f"- {nome} {dt:%d/%m}: {CODES.get(daily['weather_code'][i], '')}, min {daily['temperature_2m_min'][i]}° max {daily['temperature_2m_max'][i]}°, pioggia {daily['precipitation_probability_max'][i]}%")
return "\n".join(out)
TOOLS = [
{"name": "meteo", "description": "Previsioni meteo per una città: ora, oggi e i prossimi giorni.",
"parameters": {"type": "object", "properties": {"citta": {"type": "string"}, "giorni": {"type": "integer", "description": "Quanti giorni (1-7, default 2)."}}, "required": ["citta"]},
"run": previsioni},
]
+94
View File
@@ -0,0 +1,94 @@
"""Musica: controlla Apple Music (o Spotify, se aperto): riproduci, pausa, brano, playlist, volume."""
from __future__ import annotations
import subprocess
import sys
import time
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import osa # noqa: E402
def _running(name: str) -> bool:
return subprocess.run(["pgrep", "-x", name], capture_output=True).returncode == 0
def _app(launch: bool = False) -> str | None:
"""App musicale in uso: Spotify se è l'unica aperta, altrimenti Musica (avviata se launch=True)."""
if _running("Spotify") and not _running("Music"):
return "Spotify"
if _running("Music"):
return "Music"
if launch:
subprocess.run(["open", "-a", "Music"], check=False, timeout=15)
for _ in range(20):
if _running("Music"):
time.sleep(2)
return "Music"
time.sleep(0.5)
return None
def controllo(params: dict, ctx: dict) -> str:
azione = str(params.get("azione", "")).lower()
app = _app(launch=azione in ("play", "riproduci"))
if app is None:
return "Nessuna app musicale aperta: di' «riproduci» per avviare Musica, oppure apri Spotify."
cmds = {"play": "play", "riproduci": "play", "pausa": "pause", "stop": "stop", "prossimo": "next track", "avanti": "next track",
"precedente": "previous track", "indietro": "previous track"}
if azione in cmds:
osa(f'tell application "{app}" to {cmds[azione]}', timeout=20)
return f"Fatto ({azione})."
if azione in ("brano", "cosa suona", "info"):
try:
out = osa(f'tell application "{app}" to return (name of current track) & " — " & (artist of current track) & " (" & (album of current track) & ")"')
stato = osa(f'tell application "{app}" to return player state as string')
return f"{'In riproduzione' if stato == 'playing' else 'In pausa'}: {out}"
except Exception:
return "Nessun brano in riproduzione."
if azione == "volume":
v = max(0, min(100, int(params.get("valore") or 50)))
osa(f'tell application "{app}" to set sound volume to {v}')
return f"Volume musica al {v}%."
return "Azione sconosciuta. Usa: riproduci, pausa, prossimo, precedente, brano, volume."
def riproduci(params: dict, ctx: dict) -> str:
q = str(params.get("nome", "")).strip()
tipo = str(params.get("tipo", "auto")).lower()
if not q:
return "Errore: serve il nome."
app = _app(launch=True)
if app == "Spotify":
return "Con Spotify posso solo controllare la riproduzione, non cercare: avvia tu il brano e poi dimmi pausa, avanti, ecc."
if tipo in ("playlist", "auto"):
try:
out = osa('on run argv\ntell application "Music"\nset p to first playlist whose name contains (item 1 of argv)\nplay p\nreturn name of p\nend tell\nend run', q)
return f"Riproduco la playlist «{out}»."
except Exception:
if tipo == "playlist":
return f"Nessuna playlist chiamata '{q}'."
for prop in ("artist", "album", "name"):
try:
out = osa(f'''on run argv
tell application "Music"
set trk to (first track of library playlist 1 whose {prop} contains (item 1 of argv))
play trk
return (name of trk) & " — " & (artist of trk)
end tell
end run''', q, timeout=60)
return f"Riproduco: {out}."
except Exception:
continue
return f"Non ho trovato '{q}' nella libreria di Musica."
TOOLS = [
{"name": "musica_controllo", "description": "Controlla la musica: riproduci, pausa, prossimo, precedente, brano (cosa sta suonando), volume (0-100).",
"parameters": {"type": "object", "properties": {"azione": {"type": "string", "enum": ["riproduci", "pausa", "prossimo", "precedente", "brano", "volume"]},
"valore": {"type": "integer", "description": "Per volume: 0-100."}}, "required": ["azione"]}, "run": controllo},
{"name": "musica_riproduci", "description": "Cerca e riproduce una playlist, un artista, un album o un brano nella libreria di Apple Music.",
"parameters": {"type": "object", "properties": {"nome": {"type": "string"}, "tipo": {"type": "string", "enum": ["auto", "playlist", "artista", "album", "brano"]}}, "required": ["nome"]},
"run": riproduci},
]
+80
View File
@@ -0,0 +1,80 @@
"""Note di macOS: crea, cerca e legge note (AppleScript)."""
from __future__ import annotations
import html
import re
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import osa # noqa: E402
SEP = "␟"
def _testo(body_html: str) -> str:
t = re.sub(r"<br\s*/?>|</div>|</p>|</li>", "\n", body_html, flags=re.I)
t = html.unescape(re.sub(r"<[^>]+>", "", t))
return "\n".join(line.strip() for line in t.splitlines() if line.strip())
def crea(params: dict, ctx: dict) -> str:
titolo = str(params.get("titolo", "")).strip() or "Nota"
testo = str(params.get("testo", "")).strip()
body = f"<h1>{html.escape(titolo)}</h1>" + "".join(f"<div>{html.escape(l) or '<br>'}</div>" for l in testo.splitlines())
osa('on run argv\ntell application "Notes" to make new note at folder "Notes" with properties {name:item 1 of argv, body:item 2 of argv}\nend run', titolo, body)
return f"Nota creata: «{titolo}»."
def cerca(params: dict, ctx: dict) -> str:
q = str(params.get("query", "")).strip()
if not q:
return "Errore: serve il campo 'query'."
out = osa(f'''on run argv
set sep to "{SEP}"
set out to ""
set n to 0
tell application "Notes"
repeat with nt in (notes whose name contains (item 1 of argv) or plaintext contains (item 1 of argv))
set n to n + 1
if n > 8 then exit repeat
set out to out & (name of nt) & sep & ((modification date of nt) as string) & linefeed
end repeat
end tell
return out
end run''', q, timeout=90)
rows = [l.split(SEP) for l in out.splitlines() if l]
if not rows:
return f"Nessuna nota trovata per '{q}'."
return f"Note trovate per '{q}':\n" + "\n".join(f"- {r[0]} ({r[1]})" for r in rows) + "\nPer leggerne una usa note_leggi con il titolo."
def leggi(params: dict, ctx: dict) -> str:
titolo = str(params.get("titolo", "")).strip()
if not titolo:
return "Errore: serve il titolo."
out = osa('on run argv\ntell application "Notes" to return body of (first note whose name contains (item 1 of argv))\nend run', titolo, timeout=60)
testo = _testo(out)
return testo[:2500] + (" […]" if len(testo) > 2500 else "") if testo else "Nota vuota."
def aggiungi(params: dict, ctx: dict) -> str:
titolo = str(params.get("titolo", "")).strip()
testo = str(params.get("testo", "")).strip()
if not titolo or not testo:
return "Errore: servono titolo (della nota esistente) e testo da aggiungere."
extra = "".join(f"<div>{html.escape(l) or '<br>'}</div>" for l in testo.splitlines())
osa('on run argv\ntell application "Notes"\nset nt to first note whose name contains (item 1 of argv)\nset body of nt to (body of nt) & (item 2 of argv)\nend tell\nend run', titolo, extra)
return f"Aggiunto alla nota «{titolo}»."
TOOLS = [
{"name": "note_crea", "description": "Crea una nuova nota nell'app Note con titolo e testo. Usalo per 'prendi nota', 'appunta che…'.",
"parameters": {"type": "object", "properties": {"titolo": {"type": "string"}, "testo": {"type": "string"}}, "required": ["testo"]}, "run": crea},
{"name": "note_aggiungi", "description": "Aggiunge testo in fondo a una nota esistente, cercata per titolo.",
"parameters": {"type": "object", "properties": {"titolo": {"type": "string"}, "testo": {"type": "string"}}, "required": ["titolo", "testo"]}, "run": aggiungi},
{"name": "note_cerca", "description": "Cerca note per parola nel titolo o nel testo.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]}, "run": cerca},
{"name": "note_leggi", "description": "Legge il contenuto di una nota cercata per titolo.",
"parameters": {"type": "object", "properties": {"titolo": {"type": "string"}}, "required": ["titolo"]}, "run": leggi},
]
+149
View File
@@ -0,0 +1,149 @@
"""Promemoria di macOS: elenca, aggiunge e completa promemoria (EventKit, sincronizzati con iPhone)."""
from __future__ import annotations
import threading
from datetime import datetime, timedelta
_store = None
_lock = threading.Lock()
def _get_store():
global _store
with _lock:
if _store is not None:
return _store
import EventKit
store = EventKit.EKEventStore.alloc().init()
done, res = threading.Event(), {}
def cb(granted, err):
res["granted"] = bool(granted)
done.set()
if hasattr(store, "requestFullAccessToRemindersWithCompletion_"):
store.requestFullAccessToRemindersWithCompletion_(cb)
else:
store.requestAccessToEntityType_completion_(EventKit.EKEntityTypeReminder, cb)
done.wait(60)
if not res.get("granted"):
raise RuntimeError("Accesso ai Promemoria negato: concedilo in Impostazioni di Sistema > Privacy e sicurezza > Promemoria.")
_store = store
return store
def _parse_dt(s: str) -> datetime | None:
s = (s or "").strip().replace("T", " ")
if not s:
return None
for fmt in ("%Y-%m-%d %H:%M", "%Y-%m-%d"):
try:
return datetime.strptime(s, fmt)
except ValueError:
continue
raise ValueError(f"Data non valida: '{s}'. Usa AAAA-MM-GG HH:MM.")
def _due(r) -> datetime | None:
comps = r.dueDateComponents()
if comps is None:
return None
try:
return datetime(comps.year(), comps.month(), comps.day(), comps.hour() if comps.hour() != 9223372036854775807 else 0,
comps.minute() if comps.minute() != 9223372036854775807 else 0)
except Exception:
return None
def _fetch_incomplete():
store = _get_store()
pred = store.predicateForIncompleteRemindersWithDueDateStarting_ending_calendars_(None, None, None)
done, out = threading.Event(), []
def cb(items):
out.extend(list(items or []))
done.set()
store.fetchRemindersMatchingPredicate_completion_(pred, cb)
done.wait(30)
return out
def _fmt(r) -> str:
d = _due(r)
when = f" (entro {d:%d/%m %H:%M})" if d and (d.hour or d.minute) else (f" (entro {d:%d/%m})" if d else "")
return f"- {r.title()}{when} [{r.calendar().title()}]"
def elenca(params: dict, ctx: dict) -> str:
items = _fetch_incomplete()
quando = str(params.get("quando", "tutti")).lower()
today = datetime.now().date()
if quando == "oggi":
items = [r for r in items if (d := _due(r)) and d.date() <= today]
elif quando == "settimana":
items = [r for r in items if (d := _due(r)) and d.date() <= today + timedelta(days=7)]
items.sort(key=lambda r: (_due(r) or datetime.max))
if not items:
return "Nessun promemoria da fare" + (" per oggi." if quando == "oggi" else ".")
return f"Promemoria da fare ({len(items)}):\n" + "\n".join(_fmt(r) for r in items[:30])
def aggiungi(params: dict, ctx: dict) -> str:
import EventKit
from Foundation import NSDateComponents
store = _get_store()
titolo = str(params.get("titolo", "")).strip()
if not titolo:
return "Errore: serve il titolo."
r = EventKit.EKReminder.reminderWithEventStore_(store)
r.setTitle_(titolo)
lista = str(params.get("lista", "")).strip()
cal = None
if lista:
for c in store.calendarsForEntityType_(EventKit.EKEntityTypeReminder):
if c.title().lower() == lista.lower():
cal = c
r.setCalendar_(cal or store.defaultCalendarForNewReminders())
scadenza = _parse_dt(str(params.get("scadenza", "")))
if scadenza:
comps = NSDateComponents.alloc().init()
comps.setYear_(scadenza.year); comps.setMonth_(scadenza.month); comps.setDay_(scadenza.day)
comps.setHour_(scadenza.hour); comps.setMinute_(scadenza.minute)
r.setDueDateComponents_(comps)
from Foundation import NSDate
alarm = EventKit.EKAlarm.alarmWithAbsoluteDate_(NSDate.dateWithTimeIntervalSince1970_(scadenza.timestamp()))
r.addAlarm_(alarm)
if params.get("note"):
r.setNotes_(str(params["note"]))
ok, err = store.saveReminder_commit_error_(r, True, None)
return f"Promemoria aggiunto: {_fmt(r)[2:]}" if ok else f"Errore: {err}"
def completa(params: dict, ctx: dict) -> str:
titolo = str(params.get("titolo", "")).strip().lower()
if not titolo:
return "Errore: serve il titolo (anche parziale)."
found = [r for r in _fetch_incomplete() if titolo in (r.title() or "").lower()]
if not found:
return "Nessun promemoria con quel titolo."
if len(found) > 1 and not params.get("tutti"):
return "Ho trovato più promemoria, dimmi quale:\n" + "\n".join(_fmt(r) for r in found)
store = _get_store()
for r in found:
r.setCompleted_(True)
store.saveReminder_commit_error_(r, True, None)
return "Completato: " + ", ".join(r.title() for r in found)
TOOLS = [
{"name": "promemoria_elenca", "description": "Elenca i promemoria da fare dell'app Promemoria (tutti, di oggi o della settimana).",
"parameters": {"type": "object", "properties": {"quando": {"type": "string", "enum": ["tutti", "oggi", "settimana"]}}}, "run": elenca},
{"name": "promemoria_aggiungi", "description": "Aggiunge un promemoria (o una voce a una lista, per esempio 'Spesa'). Converti 'domani alle 17' in AAAA-MM-GG HH:MM usando la data del prompt.",
"parameters": {"type": "object", "properties": {"titolo": {"type": "string"}, "scadenza": {"type": "string", "description": "AAAA-MM-GG HH:MM, facoltativa."},
"lista": {"type": "string", "description": "Nome della lista (facoltativo)."}, "note": {"type": "string"}}, "required": ["titolo"]},
"run": aggiungi},
{"name": "promemoria_completa", "description": "Segna come fatto un promemoria cercandolo per titolo.",
"parameters": {"type": "object", "properties": {"titolo": {"type": "string"}, "tutti": {"type": "boolean", "description": "Completa tutti quelli che corrispondono."}}, "required": ["titolo"]},
"run": completa},
]
+155
View File
@@ -0,0 +1,155 @@
"""Telegram con il tuo account: chat non lette, lettura, ricerca e invio (con conferma)."""
from __future__ import annotations
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import confirm_or_param, CONFERMATO # noqa: E402
from avatar import telegram_client as tg # noqa: E402
def _name(entity) -> str:
for attr in ("title",):
if getattr(entity, attr, None):
return entity.title
return f"{getattr(entity, 'first_name', '') or ''} {getattr(entity, 'last_name', '') or ''}".strip() or getattr(entity, "username", "") or "?"
def _preview(msg) -> str:
t = (msg.message or "").strip() if msg else ""
if not t and msg is not None:
t = "[media]" if getattr(msg, "media", None) else ""
return " ".join(t.split())[:140]
def non_letti(params: dict, ctx: dict) -> str:
lim = max(1, min(int(params.get("limite") or 10), 30))
async def go(client):
out, total = [], 0
async for d in client.iter_dialogs(limit=100):
if d.unread_count > 0 and not d.is_channel or (d.is_channel and d.is_group and d.unread_count > 0):
total += d.unread_count
sender = ""
if d.message and d.is_group:
try:
s = await d.message.get_sender()
sender = _name(s) + ": "
except Exception:
pass
out.append(f"- {d.name} ({d.unread_count} non lett{'o' if d.unread_count == 1 else 'i'}): {sender}{_preview(d.message)}")
if len(out) >= lim:
break
return out, total
out, total = tg.run(go)
if not out:
return "Nessun messaggio Telegram non letto (canali esclusi)."
return f"Telegram, {total} messaggi non letti:\n" + "\n".join(out) + "\nPer leggere una chat usa telegram_leggi con il nome."
async def _find_dialog(client, name: str):
name_l = name.lower()
best = None
async for d in client.iter_dialogs(limit=300):
n = (d.name or "").lower()
if n == name_l:
return d
if name_l in n and best is None:
best = d
if best is None:
try:
ent = await client.get_entity(name)
return ent
except Exception:
return None
return best
def leggi(params: dict, ctx: dict) -> str:
chat = str(params.get("chat", "")).strip()
n = max(1, min(int(params.get("numero") or 10), 40))
if not chat:
return "Errore: serve il nome della chat o del contatto."
async def go(client):
d = await _find_dialog(client, chat)
if d is None:
return None
entity = getattr(d, "entity", d)
lines = []
async for m in client.iter_messages(entity, limit=n):
who = "tu" if m.out else _name(await m.get_sender()) if m.sender_id else "?"
lines.append(f"[{m.date.astimezone():%d/%m %H:%M}] {who}: {_preview(m)}")
try:
await client.send_read_acknowledge(entity)
except Exception:
pass
return getattr(d, "name", None) or _name(entity), list(reversed(lines))
res = tg.run(go)
if res is None:
return f"Nessuna chat trovata per '{chat}'."
name, lines = res
return f"Chat «{name}», ultimi messaggi:\n" + "\n".join(lines)
def cerca(params: dict, ctx: dict) -> str:
q = str(params.get("query", "")).strip()
if not q:
return "Errore: serve la query."
async def go(client):
lines = []
async for m in client.iter_messages(None, search=q, limit=15):
try:
chat = _name(await m.get_chat())
except Exception:
chat = "?"
lines.append(f"[{m.date.astimezone():%d/%m %H:%M}] {chat}: {_preview(m)}")
return lines
lines = tg.run(go)
return f"Messaggi Telegram con '{q}':\n" + "\n".join(lines) if lines else f"Nessun messaggio trovato per '{q}'."
def invia(params: dict, ctx: dict) -> str:
chat = str(params.get("chat", "")).strip()
testo = str(params.get("testo", "")).strip()
if not chat or not testo:
return "Errore: servono destinatario (nome della chat, @username o numero) e testo."
async def resolve(client):
d = await _find_dialog(client, chat)
if d is None:
return None
return getattr(d, "name", None) or _name(d)
name = tg.run(resolve)
if name is None:
return f"Non trovo '{chat}' tra le tue chat Telegram."
def do() -> str:
async def send(client):
d = await _find_dialog(client, chat)
await client.send_message(getattr(d, "entity", d), testo)
return f"Messaggio Telegram inviato a {name}."
msg = tg.run(send)
if ctx.get("say"):
ctx["say"](msg)
return msg
return confirm_or_param(ctx, params, "telegram_invia", "Inviare su Telegram?", f"A: {name}\n{testo[:200]}", do)
TOOLS = [
{"name": "telegram_non_letti", "description": "Elenca le chat Telegram con messaggi non letti e l'ultimo messaggio di ciascuna.",
"parameters": {"type": "object", "properties": {"limite": {"type": "integer"}}}, "run": non_letti},
{"name": "telegram_leggi", "description": "Legge gli ultimi messaggi di una chat o di un contatto Telegram (per nome) e li segna come letti.",
"parameters": {"type": "object", "properties": {"chat": {"type": "string"}, "numero": {"type": "integer"}}, "required": ["chat"]}, "run": leggi},
{"name": "telegram_cerca", "description": "Cerca messaggi in tutte le chat Telegram per parola.",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]}, "run": cerca},
{"name": "telegram_invia", "description": "Invia un messaggio Telegram a un contatto o a un gruppo (per nome o @username). Azione irreversibile con conferma sullo schermo; con [CONFIRMATION_PENDING] di' all'utente di confermare e non dire che è inviato.",
"parameters": {"type": "object", "properties": {"chat": {"type": "string"}, "testo": {"type": "string"}, "confermato": CONFERMATO}, "required": ["chat", "testo"]}, "run": invia},
]
+109
View File
@@ -0,0 +1,109 @@
"""Timer, sveglie e attività programmate: l'app avvisa a voce o esegue un comando all'ora stabilita, anche ogni giorno."""
from __future__ import annotations
import json
import sys
import uuid
from datetime import datetime, timedelta
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.settings import DATA_DIR # noqa: E402
FILE = DATA_DIR / "timers.json"
def load() -> list[dict]:
try:
return json.loads(FILE.read_text(encoding="utf-8"))
except Exception:
return []
def save(items: list[dict]) -> None:
FILE.parent.mkdir(parents=True, exist_ok=True)
FILE.write_text(json.dumps(items, ensure_ascii=False, indent=2), encoding="utf-8")
def _when(params: dict) -> datetime:
now = datetime.now()
if params.get("minuti"):
return now + timedelta(minutes=float(params["minuti"]))
s = str(params.get("ora", "")).strip().replace("T", " ")
if not s:
raise ValueError("Serve 'minuti' oppure 'ora'.")
if len(s) <= 5: # HH:MM oggi (o domani se già passata)
t = datetime.strptime(s, "%H:%M").time()
dt = datetime.combine(now.date(), t)
return dt if dt > now else dt + timedelta(days=1)
return datetime.strptime(s, "%Y-%m-%d %H:%M")
def imposta(params: dict, ctx: dict) -> str:
testo = str(params.get("testo", "")).strip() or "Timer scaduto"
try:
at = _when(params)
except ValueError as err:
return f"Errore: {err}"
item = {"id": uuid.uuid4().hex[:6], "at": at.strftime("%Y-%m-%d %H:%M"), "testo": testo,
"ripeti": "giorno" if params.get("ogni_giorno") else "no", "comando": bool(params.get("esegui_come_comando"))}
items = load(); items.append(item); save(items)
quando = at.strftime("%H:%M") if at.date() == datetime.now().date() else at.strftime("%d/%m alle %H:%M")
tipo = "comando" if item["comando"] else "avviso"
return f"{tipo.capitalize()} programmato per le {quando}{' ogni giorno' if item['ripeti'] == 'giorno' else ''}: «{testo}» [id {item['id']}]."
def elenca(params: dict, ctx: dict) -> str:
items = sorted(load(), key=lambda x: x["at"])
if not items:
return "Nessun timer o attività programmata."
return "Programmati:\n" + "\n".join(f"- [{i['id']}] {i['at']}{' ogni giorno' if i.get('ripeti') == 'giorno' else ''}: {i['testo']}{' (comando)' if i.get('comando') else ''}" for i in items)
def annulla(params: dict, ctx: dict) -> str:
key = str(params.get("id", "")).strip().lower()
items = load()
keep = [i for i in items if i["id"] != key and key not in i["testo"].lower()] if key else []
if key and len(keep) == len(items):
return "Nessun timer corrispondente."
if not key:
save([]); return "Tutti i timer annullati."
save(keep)
return f"Annullati {len(items) - len(keep)} timer."
def due(now: datetime | None = None) -> list[dict]:
"""Restituisce gli elementi scaduti e aggiorna il file (usato dallo scheduler dell'app)."""
now = now or datetime.now()
items, fired, keep = load(), [], []
for i in items:
try:
at = datetime.strptime(i["at"], "%Y-%m-%d %H:%M")
except ValueError:
continue
if at <= now:
fired.append(i)
if i.get("ripeti") == "giorno":
nxt = at + timedelta(days=1)
while nxt <= now:
nxt += timedelta(days=1)
keep.append({**i, "at": nxt.strftime("%Y-%m-%d %H:%M")})
else:
keep.append(i)
if fired:
save(keep)
return fired
TOOLS = [
{"name": "timer_imposta",
"description": "Imposta un timer ('tra 10 minuti'), una sveglia/avviso a un'ora, o un'attività programmata anche giornaliera. Con esegui_come_comando=true, all'ora stabilita il testo viene eseguito come richiesta all'assistente (es. 'leggimi la posta'), altrimenti viene annunciato a voce.",
"parameters": {"type": "object", "properties": {"testo": {"type": "string", "description": "Cosa annunciare, o il comando da eseguire."},
"minuti": {"type": "number", "description": "Fra quanti minuti."},
"ora": {"type": "string", "description": "HH:MM oppure AAAA-MM-GG HH:MM."},
"ogni_giorno": {"type": "boolean"}, "esegui_come_comando": {"type": "boolean"}}, "required": ["testo"]},
"run": imposta},
{"name": "timer_elenca", "description": "Elenca timer e attività programmate.", "parameters": {"type": "object", "properties": {}}, "run": elenca},
{"name": "timer_annulla", "description": "Annulla un timer per id o parola del testo; senza id annulla tutti.",
"parameters": {"type": "object", "properties": {"id": {"type": "string"}}}, "run": annulla},
]
+279
View File
@@ -0,0 +1,279 @@
"""Archivio delle chat WhatsApp esportate (.txt o .zip): indice locale, ricerca, lettura per periodo, riepiloghi."""
from __future__ import annotations
import re
import shutil
import sqlite3
import sys
import zipfile
from datetime import datetime, timedelta
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.settings import DATA_DIR # noqa: E402
ARCHIVE = DATA_DIR / "whatsapp"
DB = ARCHIVE / "index.sqlite"
MAX_CHARS = 12000
# iOS: [21/11/20, 13:56:41] Nome: testo — Android: 21/11/20, 13:56 - Nome: testo
RE_IOS = re.compile(r"^‎?\[(\d{1,2}/\d{1,2}/\d{2,4}),? (\d{1,2}:\d{2}(?::\d{2})?)\] ([^:]{1,80}?): (.*)$")
RE_AND = re.compile(r"^‎?(\d{1,2}/\d{1,2}/\d{2,4}),? (\d{1,2}:\d{2}(?::\d{2})?) - ([^:]{1,80}?): (.*)$")
SKIP = ("crittografati end-to-end", "end-to-end encrypted")
MEDIA = {"immagine omessa": "[immagine]", "video omesso": "[video]", "audio omesso": "[audio]", "sticker omesso": "[sticker]",
"GIF omessa": "[gif]", "documento omesso": "[documento]", "image omitted": "[immagine]", "video omitted": "[video]"}
def _parse_ts(d: str, t: str) -> datetime | None:
for fmt in ("%d/%m/%y %H:%M:%S", "%d/%m/%Y %H:%M:%S", "%d/%m/%y %H:%M", "%d/%m/%Y %H:%M"):
try:
return datetime.strptime(f"{d} {t}", fmt)
except ValueError:
continue
return None
def parse(text: str):
msgs = []
for raw in text.splitlines():
line = raw.rstrip("\r").replace("‎", "")
m = RE_IOS.match(line) or RE_AND.match(line)
if m:
ts = _parse_ts(m.group(1), m.group(2))
body = m.group(4).strip()
if ts is None or any(s in body for s in SKIP):
continue
for k, v in MEDIA.items():
if body == k:
body = v
msgs.append([ts, m.group(3).strip(), body])
elif msgs and line.strip():
msgs[-1][2] += "\n" + line.strip()
return msgs
def _db() -> sqlite3.Connection:
ARCHIVE.mkdir(parents=True, exist_ok=True)
con = sqlite3.connect(DB)
con.execute("CREATE TABLE IF NOT EXISTS chats (name TEXT PRIMARY KEY, file TEXT, mtime REAL, count INTEGER, first TEXT, last TEXT, participants TEXT)")
con.execute("CREATE TABLE IF NOT EXISTS messages (id INTEGER PRIMARY KEY, chat TEXT, ts TEXT, sender TEXT, text TEXT)")
con.execute("CREATE INDEX IF NOT EXISTS idx_chat_ts ON messages(chat, ts)")
con.execute("CREATE VIRTUAL TABLE IF NOT EXISTS fts USING fts5(text, content='messages', content_rowid='id', tokenize='unicode61 remove_diacritics 2')")
con.execute("CREATE TRIGGER IF NOT EXISTS messages_ai AFTER INSERT ON messages BEGIN INSERT INTO fts(rowid, text) VALUES (new.id, new.text); END")
con.execute("CREATE TRIGGER IF NOT EXISTS messages_ad AFTER DELETE ON messages BEGIN INSERT INTO fts(fts, rowid, text) VALUES ('delete', old.id, old.text); END")
for col in ("source TEXT", "jid TEXT", "from_me INTEGER DEFAULT 0"):
try:
con.execute(f"ALTER TABLE messages ADD COLUMN {col}")
except sqlite3.OperationalError:
pass
con.execute("PRAGMA journal_mode=WAL")
return con
def _index_file(con: sqlite3.Connection, name: str, path: Path) -> int:
msgs = parse(path.read_text(encoding="utf-8", errors="replace"))
con.execute("DELETE FROM messages WHERE chat = ? AND (source IS NULL OR source = 'export')", (name,))
con.executemany("INSERT INTO messages (chat, ts, sender, text, source) VALUES (?, ?, ?, ?, 'export')",
[(name, ts.strftime("%Y-%m-%d %H:%M:%S"), s, t) for ts, s, t in msgs])
con.execute("INSERT INTO fts(fts) VALUES ('rebuild')")
senders = sorted({s for _, s, _ in msgs}, key=lambda s: -sum(1 for _, x, _ in msgs if x == s))
con.execute("INSERT OR REPLACE INTO chats VALUES (?, ?, ?, ?, ?, ?, ?)",
(name, str(path), path.stat().st_mtime, len(msgs), msgs[0][0].strftime("%Y-%m-%d") if msgs else "",
msgs[-1][0].strftime("%Y-%m-%d") if msgs else "", ", ".join(senders[:6])))
con.commit()
return len(msgs)
def _sync(con: sqlite3.Connection) -> list[str]:
"""Indicizza i .txt nuovi o modificati presenti nella cartella archivio."""
ARCHIVE.mkdir(parents=True, exist_ok=True)
known = {r[0]: r for r in con.execute("SELECT name, file, mtime FROM chats")}
done = []
for path in sorted(ARCHIVE.glob("*.txt")):
name = path.stem
if name in known and abs(known[name][2] - path.stat().st_mtime) < 1:
continue
n = _index_file(con, name, path)
done.append(f"{name} ({n} messaggi)")
return done
def _chat_name(con: sqlite3.Connection, chat: str) -> str | None:
names = [r[0] for r in con.execute("SELECT name FROM chats")]
if not chat:
return names[0] if len(names) == 1 else None
q = chat.lower()
return next((n for n in names if n.lower() == q), None) or next((n for n in names if q in n.lower()), None)
def _range(params: dict) -> tuple[str | None, str | None]:
da, a = str(params.get("da", "")).strip(), str(params.get("a", "")).strip()
today = datetime.now().date()
alias = {"oggi": (today, today), "ieri": (today - timedelta(days=1), today - timedelta(days=1)),
"settimana": (today - timedelta(days=7), today), "mese": (today - timedelta(days=30), today)}
if da.lower() in alias:
d1, d2 = alias[da.lower()]
return d1.isoformat(), d2.isoformat() + " 23:59:59"
return (da or None), ((a[:10] + " 23:59:59") if a else None)
def _fmt(rows) -> str:
return "\n".join(f"[{ts[:16]}] {s}: {t}" for ts, s, t in rows)
def importa(params: dict, ctx: dict) -> str:
src = Path(str(params.get("percorso", ""))).expanduser()
nome = str(params.get("nome", "")).strip()
if not src.exists():
return f"Percorso non trovato: {src}"
ARCHIVE.mkdir(parents=True, exist_ok=True)
txt: Path | None = None
if src.is_dir():
cands = list(src.glob("*.txt"))
txt = next((c for c in cands if c.name == "_chat.txt"), cands[0] if cands else None)
elif src.suffix.lower() == ".zip":
with zipfile.ZipFile(src) as z:
members = [m for m in z.namelist() if m.endswith(".txt") and not m.startswith("__MACOSX") and not Path(m).name.startswith(".")]
member = next((m for m in members if Path(m).name == "_chat.txt"), members[0] if members else None)
if not member:
return "Nello zip non c'è un file .txt."
tmp = ARCHIVE / "_import.txt"
tmp.write_bytes(z.read(member))
txt = tmp
else:
txt = src
if txt is None:
return "Nessun file .txt trovato."
con = _db()
if not nome:
# Gli zip di WhatsApp si chiamano "WhatsApp Chat - Nome.zip": usa quel nome.
m = re.match(r"^WhatsApp Chat - (.+)$", src.stem, flags=re.I)
if m:
nome = m.group(1).strip()
if not nome:
msgs = parse(txt.read_text(encoding="utf-8", errors="replace"))
senders = {}
for _, s, _ in msgs:
senders[s] = senders.get(s, 0) + 1
top = sorted(senders, key=senders.get, reverse=True)
nome = " & ".join(top[:2]) if len(top) > 1 else (top[0] if top else src.stem)
safe = re.sub(r'[\\/:*?"<>|]+', "_", nome).strip() or "chat"
dest = ARCHIVE / f"{safe}.txt"
if txt.resolve() != dest.resolve():
shutil.copyfile(txt, dest)
if (ARCHIVE / "_import.txt").exists():
(ARCHIVE / "_import.txt").unlink()
n = _index_file(con, safe, dest)
con.close()
return f"Chat importata come «{safe}»: {n} messaggi indicizzati."
def importa_cartella(params: dict, ctx: dict) -> str:
folder = Path(str(params.get("percorso", ""))).expanduser()
if not folder.is_dir():
return f"Cartella non trovata: {folder}"
files = sorted([p for p in folder.iterdir() if p.suffix.lower() in (".zip", ".txt") and not p.name.startswith(".")])
if not files:
return "Nessun file .zip o .txt nella cartella."
results, errors = [], []
for f in files:
try:
results.append(importa({"percorso": str(f)}, ctx))
except Exception as err:
errors.append(f"{f.name}: {err}")
out = f"Importate {len(results)} chat:\n" + "\n".join("- " + r.replace("Chat importata come ", "") for r in results)
if errors:
out += "\nErrori:\n" + "\n".join("- " + e for e in errors)
return out
def elenca(params: dict, ctx: dict) -> str:
con = _db()
_sync(con)
rows = con.execute("SELECT name, count, first, last, participants FROM chats ORDER BY name").fetchall()
con.close()
if not rows:
return f"Archivio vuoto. Esporta una chat da WhatsApp e importala con whatsapp_importa, oppure copia il .txt in {ARCHIVE}."
return "Chat in archivio:\n" + "\n".join(f"- {n}: {c} messaggi dal {f} al {l} (partecipanti: {p})" for n, c, f, l, p in rows)
def cerca(params: dict, ctx: dict) -> str:
q = str(params.get("query", "")).strip()
if not q:
return "Errore: serve la query."
lim = max(1, min(int(params.get("limite") or 15), 60))
con = _db()
_sync(con)
chat = _chat_name(con, str(params.get("chat", "")))
da, a = _range(params)
sql = "SELECT m.ts, m.sender, m.text FROM fts JOIN messages m ON m.id = fts.rowid WHERE fts MATCH ?"
args: list = ['"' + q.replace('"', '') + '"*' if " " not in q else " ".join('"' + w.replace('"', '') + '"' for w in q.split())]
if chat:
sql += " AND m.chat = ?"; args.append(chat)
if da:
sql += " AND m.ts >= ?"; args.append(da)
if a:
sql += " AND m.ts <= ?"; args.append(a)
sql += " ORDER BY m.ts DESC LIMIT ?"; args.append(lim)
rows = con.execute(sql, args).fetchall()
con.close()
if not rows:
return f"Nessun messaggio contiene '{q}'."
return f"Messaggi con '{q}' ({len(rows)}, i più recenti):\n" + _fmt(reversed(rows))
def leggi(params: dict, ctx: dict) -> str:
con = _db()
_sync(con)
chat = _chat_name(con, str(params.get("chat", "")))
if chat is None:
con.close()
return "Chat non trovata: usa whatsapp_chat_elenca per vedere i nomi."
da, a = _range(params)
n = max(1, min(int(params.get("ultimi") or 0) or 0, 400))
if da or a:
sql, args = "SELECT ts, sender, text FROM messages WHERE chat = ?", [chat]
if da:
sql += " AND ts >= ?"; args.append(da)
if a:
sql += " AND ts <= ?"; args.append(a)
rows = con.execute(sql + " ORDER BY ts LIMIT 400", args).fetchall()
else:
rows = list(reversed(con.execute("SELECT ts, sender, text FROM messages WHERE chat = ? ORDER BY ts DESC LIMIT ?", (chat, n or 30)).fetchall()))
con.close()
if not rows:
return "Nessun messaggio nel periodo."
text = _fmt(rows)
if len(text) > MAX_CHARS:
text = text[-MAX_CHARS:]
text = "[…]\n" + text[text.find("\n") + 1:]
return f"Chat «{chat}», {len(rows)} messaggi:\n{text}"
def statistiche(params: dict, ctx: dict) -> str:
con = _db()
_sync(con)
chat = _chat_name(con, str(params.get("chat", "")))
if chat is None:
con.close()
return "Chat non trovata."
rows = con.execute("SELECT sender, COUNT(*) FROM messages WHERE chat = ? GROUP BY sender ORDER BY 2 DESC", (chat,)).fetchall()
months = con.execute("SELECT substr(ts,1,7), COUNT(*) FROM messages WHERE chat = ? GROUP BY 1 ORDER BY 1 DESC LIMIT 12", (chat,)).fetchall()
con.close()
return (f"Chat «{chat}»: " + ", ".join(f"{s} {c} messaggi" for s, c in rows) + ".\nUltimi mesi: " +
", ".join(f"{m} {c}" for m, c in months))
TOOLS = [
{"name": "whatsapp_importa", "description": "Importa una chat WhatsApp esportata (.txt, .zip o cartella con _chat.txt) nell'archivio locale e la indicizza. Il nome è facoltativo: altrimenti usa i partecipanti.",
"parameters": {"type": "object", "properties": {"percorso": {"type": "string"}, "nome": {"type": "string"}}, "required": ["percorso"]}, "run": importa},
{"name": "whatsapp_importa_cartella", "description": "Importa tutte le chat WhatsApp esportate (.zip o .txt) presenti in una cartella.",
"parameters": {"type": "object", "properties": {"percorso": {"type": "string"}}, "required": ["percorso"]}, "run": importa_cartella},
{"name": "whatsapp_chat_elenca", "description": "Elenca le chat WhatsApp in archivio con numero di messaggi, periodo e partecipanti.",
"parameters": {"type": "object", "properties": {}}, "run": elenca},
{"name": "whatsapp_cerca", "description": "Cerca parole nei messaggi WhatsApp archiviati (tutte le chat o una sola), con periodo facoltativo (da/a AAAA-MM-GG oppure da='settimana'|'mese').",
"parameters": {"type": "object", "properties": {"query": {"type": "string"}, "chat": {"type": "string"}, "da": {"type": "string"}, "a": {"type": "string"}, "limite": {"type": "integer"}}, "required": ["query"]}, "run": cerca},
{"name": "whatsapp_leggi", "description": "Legge i messaggi di una chat: gli ultimi N oppure un periodo (da/a). Usalo per riassumere o rispondere a domande su cosa è stato detto.",
"parameters": {"type": "object", "properties": {"chat": {"type": "string"}, "ultimi": {"type": "integer"}, "da": {"type": "string"}, "a": {"type": "string"}}}, "run": leggi},
{"name": "whatsapp_statistiche", "description": "Quanti messaggi per persona e per mese in una chat.",
"parameters": {"type": "object", "properties": {"chat": {"type": "string"}}}, "run": statistiche},
]
+76
View File
@@ -0,0 +1,76 @@
"""WhatsApp in tempo reale (dispositivo collegato): novità, stato e invio con conferma."""
from __future__ import annotations
import sqlite3
import sys
from datetime import datetime, timedelta
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from avatar.macos import confirm_or_param, CONFERMATO # noqa: E402
from avatar import whatsapp_bridge as wb # noqa: E402
from avatar.settings import DATA_DIR # noqa: E402
DB = DATA_DIR / "whatsapp" / "index.sqlite"
def stato(params: dict, ctx: dict) -> str:
return wb.status_text()
def novita(params: dict, ctx: dict) -> str:
minuti = int(params.get("minuti") or 0)
lim = max(1, min(int(params.get("limite") or 20), 60))
if not DB.exists():
return "Nessun messaggio ricevuto finora."
since = (datetime.now() - timedelta(minutes=minuti)).strftime("%Y-%m-%d %H:%M:%S") if minuti else (datetime.now().strftime("%Y-%m-%d") + " 00:00:00")
con = sqlite3.connect(f"file:{DB}?mode=ro", uri=True)
try:
rows = con.execute("SELECT ts, chat, sender, text FROM messages WHERE source = 'live' AND from_me = 0 AND ts >= ? ORDER BY ts DESC LIMIT ?", (since, lim)).fetchall()
finally:
con.close()
if not rows:
return "Nessun messaggio WhatsApp nuovo" + (f" negli ultimi {minuti} minuti." if minuti else " oggi.") + (" " + wb.status_text() if not wb.running() else "")
lines = [f"[{ts[11:16]}] {chat}{' — ' + sender if sender != chat else ''}: {text[:200]}" for ts, chat, sender, text in reversed(rows)]
return f"Messaggi WhatsApp ricevuti ({len(rows)}):\n" + "\n".join(lines)
def invia(params: dict, ctx: dict) -> str:
a = str(params.get("a", "")).strip()
testo = str(params.get("testo", "")).strip()
if not a or not testo:
return "Errore: servono destinatario (nome del contatto o del gruppo, oppure numero) e testo."
if not wb.running():
wb.start()
try:
res = wb.resolve(a)
except Exception as err:
return f"WhatsApp non disponibile: {err}"
if not res.get("jid"):
return f"Non trovo '{a}' tra i contatti WhatsApp: prova con il numero di telefono."
name = res.get("name") or a
def do() -> str:
out = wb.send(a, testo)
msg = f"Messaggio WhatsApp inviato a {out.get('name') or name}."
if ctx.get("say"):
ctx["say"](msg)
return msg
return confirm_or_param(ctx, params, "whatsapp_invia", "Inviare su WhatsApp?", f"A: {name}\n{testo[:200]}", do)
def aggiorna_contatti(params: dict, ctx: dict) -> str:
n = wb.sync_contacts()
return f"Rubrica sincronizzata: {n} numeri associati ai nomi."
TOOLS = [
{"name": "whatsapp_aggiorna_contatti", "description": "Associa i numeri WhatsApp ai nomi della rubrica del Mac (da usare se le chat compaiono come numeri).",
"parameters": {"type": "object", "properties": {}}, "run": aggiorna_contatti},
{"name": "whatsapp_novita", "description": "Messaggi WhatsApp ricevuti in tempo reale oggi o negli ultimi N minuti (tutte le chat). Usalo per 'ci sono novità su WhatsApp?', 'chi mi ha scritto?'.",
"parameters": {"type": "object", "properties": {"minuti": {"type": "integer"}, "limite": {"type": "integer"}}}, "run": novita},
{"name": "whatsapp_invia", "description": "Invia un messaggio WhatsApp a un contatto, a un gruppo o a un numero. Azione irreversibile con conferma sullo schermo; con [CONFIRMATION_PENDING] di' all'utente di confermare e non dire che è inviato.",
"parameters": {"type": "object", "properties": {"a": {"type": "string"}, "testo": {"type": "string"}, "confermato": CONFERMATO}, "required": ["a", "testo"]}, "run": invia},
{"name": "whatsapp_stato", "description": "Stato del collegamento WhatsApp in tempo reale.", "parameters": {"type": "object", "properties": {}}, "run": stato},
]
+29
View File
@@ -0,0 +1,29 @@
[project]
name = "avatarpy"
version = "0.1.0"
description = "Add your description here"
readme = "README.md"
requires-python = ">=3.12"
dependencies = [
"anthropic>=1.7.0",
"faster-whisper>=1.2.1",
"keyring>=25.7.0",
"kokoro>=0.9.4",
"mlx-whisper>=0.4.3",
"numpy>=2.5.3",
"openai>=3.17.0",
"psutil>=7.2.2",
"pyobjc-framework-contacts>=12.2.2",
"pyobjc-framework-eventkit>=12.2.2",
"pyobjc-framework-quartz>=12.2.2",
"pyobjc-framework-vision>=12.2.2",
"pypdf>=6.19.0",
"pyqt6>=6.11.0",
"pyqt6-webengine>=6.11.0",
"python-docx>=1.2.0",
"qrcode[pil]>=8.2",
"requests>=2.34.2",
"sounddevice>=0.5.6",
"soundfile>=0.14.0",
"telethon>=1.45.0",
]
+5489
View File
File diff suppressed because it is too large. Load diff
Generated
+3577
View File
File diff suppressed because it is too large. Load diff
+141
View File
@@ -0,0 +1,141 @@
// Ponte WhatsApp per AvatarPy: si collega come "dispositivo collegato" (Baileys),
// scrive i messaggi nel database dell'archivio e offre una piccola API HTTP locale.
import http from 'node:http';
import path from 'node:path';
import fs from 'node:fs';
import { DatabaseSync } from 'node:sqlite';
import pino from 'pino';
import { Boom } from '@hapi/boom';
import makeWASocket, { useMultiFileAuthState, fetchLatestBaileysVersion, DisconnectReason, isJidGroup, jidNormalizedUser } from '@whiskeysockets/baileys';
const args = Object.fromEntries(process.argv.slice(2).map((a, i, arr) => a.startsWith('--') ? [a.slice(2), arr[i + 1]] : []).filter(Boolean));
const DATA = args.data || path.join(process.cwd(), 'data', 'whatsapp');
const PORT = Number(args.port || 8790);
const ME_NAME = args.me || 'io';
const AUTH_DIR = path.join(DATA, 'auth');
const DB_PATH = path.join(DATA, 'index.sqlite');
fs.mkdirSync(AUTH_DIR, { recursive: true });
// ── Database (stesso dell'archivio delle chat esportate) ─────────────────────
const db = new DatabaseSync(DB_PATH);
db.exec(`PRAGMA journal_mode=WAL;
CREATE TABLE IF NOT EXISTS chats (name TEXT PRIMARY KEY, file TEXT, mtime REAL, count INTEGER, first TEXT, last TEXT, participants TEXT);
CREATE TABLE IF NOT EXISTS messages (id INTEGER PRIMARY KEY, chat TEXT, ts TEXT, sender TEXT, text TEXT);
CREATE INDEX IF NOT EXISTS idx_chat_ts ON messages(chat, ts);
CREATE VIRTUAL TABLE IF NOT EXISTS fts USING fts5(text, content='messages', content_rowid='id', tokenize='unicode61 remove_diacritics 2');
CREATE TRIGGER IF NOT EXISTS messages_ai AFTER INSERT ON messages BEGIN INSERT INTO fts(rowid, text) VALUES (new.id, new.text); END;
CREATE TRIGGER IF NOT EXISTS messages_ad AFTER DELETE ON messages BEGIN INSERT INTO fts(fts, rowid, text) VALUES ('delete', old.id, old.text); END;
CREATE TABLE IF NOT EXISTS wa_contacts (jid TEXT PRIMARY KEY, name TEXT, notify TEXT);
CREATE TABLE IF NOT EXISTS wa_seen (id TEXT PRIMARY KEY);`);
for (const col of ['source TEXT', 'jid TEXT', 'from_me INTEGER DEFAULT 0']) {
try { db.exec(`ALTER TABLE messages ADD COLUMN ${col}`); } catch { /* già presente */ }
}
const stInsert = db.prepare('INSERT INTO messages (chat, ts, sender, text, source, jid, from_me) VALUES (?, ?, ?, ?, ?, ?, ?)');
const stSeen = db.prepare('INSERT OR IGNORE INTO wa_seen (id) VALUES (?)');
const stContact = db.prepare('INSERT INTO wa_contacts (jid, name, notify) VALUES (?, ?, ?) ON CONFLICT(jid) DO UPDATE SET name=COALESCE(excluded.name, name), notify=COALESCE(excluded.notify, notify)');
const stChatTouch = db.prepare(`INSERT INTO chats (name, file, mtime, count, first, last, participants) VALUES (?, 'live', 0, 1, ?, ?, ?)
ON CONFLICT(name) DO UPDATE SET count = count + 1, last = excluded.last, first = COALESCE(first, excluded.first)`);
const groupNames = new Map();
let sock = null, state = { connection: 'closed', qr: null, me: null, error: null, received: 0 };
function fmtTs(sec) { const d = new Date(Number(sec) * 1000); const p = (n) => String(n).padStart(2, '0'); return `${d.getFullYear()}-${p(d.getMonth() + 1)}-${p(d.getDate())} ${p(d.getHours())}:${p(d.getMinutes())}:${p(d.getSeconds())}`; }
function textOf(m) {
const x = m.message || {};
const inner = x.ephemeralMessage?.message || x.viewOnceMessage?.message || x;
return inner.conversation || inner.extendedTextMessage?.text || inner.imageMessage?.caption || (inner.imageMessage && '[immagine]')
|| inner.videoMessage?.caption || (inner.videoMessage && '[video]') || (inner.audioMessage && '[audio]') || (inner.documentMessage && `[documento ${inner.documentMessage.fileName || ''}]`)
|| (inner.stickerMessage && '[sticker]') || (inner.locationMessage && '[posizione]') || (inner.contactMessage && '[contatto]') || inner.reactionMessage?.text && `[reazione ${inner.reactionMessage.text}]` || '';
}
function contactName(jid) {
if (!jid) return '?';
const row = db.prepare('SELECT name, notify FROM wa_contacts WHERE jid = ?').get(jid);
return row?.name || row?.notify || jid.split('@')[0];
}
async function chatName(jid) {
if (isJidGroup(jid)) {
if (!groupNames.has(jid)) {
try { const md = await sock.groupMetadata(jid); groupNames.set(jid, md.subject); } catch { groupNames.set(jid, 'Gruppo ' + jid.split('@')[0]); }
}
return groupNames.get(jid);
}
const mine = jidNormalizedUser(sock?.user?.id || '');
if (mine && jidNormalizedUser(jid) === mine) return 'Io (note personali)';
return contactName(jid);
}
async function store(m, source) {
const id = m.key?.id, jid = m.key?.remoteJid;
if (!id || !jid || jid === 'status@broadcast') return false;
if (stSeen.run(id).changes === 0) return false;
const text = textOf(m);
if (!text) return false;
if (m.pushName && !m.key.fromMe) stContact.run(isJidGroup(jid) ? (m.key.participant || jid) : jid, null, m.pushName);
const chat = await chatName(jid);
const sender = m.key.fromMe ? (state.me?.name || ME_NAME) : (isJidGroup(jid) ? contactName(m.key.participant) : chat);
const ts = fmtTs(m.messageTimestamp);
stInsert.run(chat, ts, sender, text, source, jid, m.key.fromMe ? 1 : 0);
stChatTouch.run(chat, ts.slice(0, 10), ts.slice(0, 10), '');
state.received++;
return true;
}
async function start() {
const { state: auth, saveCreds } = await useMultiFileAuthState(AUTH_DIR);
const { version } = await fetchLatestBaileysVersion();
sock = makeWASocket({ version, auth, logger: pino({ level: 'silent' }), browser: ['AvatarPy', 'Desktop', '1.0.0'], syncFullHistory: false, markOnlineOnConnect: false });
sock.ev.on('creds.update', saveCreds);
sock.ev.on('connection.update', (u) => {
if (u.qr) { state.qr = u.qr; state.connection = 'qr'; }
if (u.connection === 'open') { state.connection = 'open'; state.qr = null; state.error = null; state.me = { id: sock.user?.id, name: sock.user?.name }; }
if (u.connection === 'close') {
const code = new Boom(u.lastDisconnect?.error)?.output?.statusCode;
state.connection = 'closed';
if (code === DisconnectReason.loggedOut) { state.error = 'Sessione chiusa da WhatsApp: ricollega con il QR.'; fs.rmSync(AUTH_DIR, { recursive: true, force: true }); fs.mkdirSync(AUTH_DIR, { recursive: true }); setTimeout(start, 2000); }
else { state.error = `Disconnesso (${code}); riconnessione…`; setTimeout(start, 3000); }
}
});
sock.ev.on('contacts.upsert', (cs) => { for (const c of cs) stContact.run(c.id, c.name || null, c.notify || null); });
sock.ev.on('contacts.update', (cs) => { for (const c of cs) if (c.id) stContact.run(c.id, c.name || null, c.notify || null); });
sock.ev.on('messaging-history.set', async ({ contacts, messages }) => {
for (const c of contacts || []) stContact.run(c.id, c.name || null, c.notify || null);
for (const m of messages || []) { try { await store(m, 'history'); } catch { /* ignora */ } }
});
sock.ev.on('messages.upsert', async ({ messages }) => {
for (const m of messages) { try { await store(m, 'live'); } catch (e) { console.error('store', e.message); } }
});
}
// ── API HTTP locale ───────────────────────────────────────────────────────────
async function resolveJid(target) {
const t = String(target || '').trim();
const digits = t.replace(/[^\d]/g, '');
if (/^\+?[\d\s]{8,}$/.test(t)) return `${digits}@s.whatsapp.net`;
const like = `%${t.toLowerCase()}%`;
const row = db.prepare('SELECT jid, name, notify FROM wa_contacts WHERE lower(name) = ? OR lower(notify) = ? OR lower(name) LIKE ? OR lower(notify) LIKE ? ORDER BY CASE WHEN lower(name) = ? THEN 0 ELSE 1 END LIMIT 1').get(t.toLowerCase(), t.toLowerCase(), like, like, t.toLowerCase());
if (row) return row.jid;
for (const [jid, name] of groupNames) if (name.toLowerCase().includes(t.toLowerCase())) return jid;
const msg = db.prepare("SELECT jid FROM messages WHERE jid IS NOT NULL AND lower(chat) LIKE ? ORDER BY ts DESC LIMIT 1").get(like);
return msg?.jid || null;
}
const server = http.createServer(async (req, res) => {
const send = (code, obj) => { res.writeHead(code, { 'Content-Type': 'application/json' }); res.end(JSON.stringify(obj)); };
try {
const url = new URL(req.url, 'http://x');
if (url.pathname === '/status') return send(200, state);
if (url.pathname === '/logout') { try { await sock?.logout(); } catch { } fs.rmSync(AUTH_DIR, { recursive: true, force: true }); state = { connection: 'closed', qr: null, me: null, error: 'Scollegato.', received: 0 }; setTimeout(start, 1000); return send(200, { ok: true }); }
if (url.pathname === '/resolve') { const jid = await resolveJid(url.searchParams.get('to')); return send(200, { jid, name: jid ? (isJidGroup(jid) ? groupNames.get(jid) : contactName(jid)) : null }); }
if (url.pathname === '/send' && req.method === 'POST') {
let body = ''; for await (const ch of req) body += ch;
const { to, text } = JSON.parse(body || '{}');
if (state.connection !== 'open') return send(409, { error: 'WhatsApp non collegato' });
const jid = await resolveJid(to);
if (!jid) return send(404, { error: 'Destinatario non trovato' });
const sent = await sock.sendMessage(jid, { text: String(text) });
try { await store(sent, 'live'); } catch { }
return send(200, { ok: true, jid, name: isJidGroup(jid) ? groupNames.get(jid) : contactName(jid) });
}
send(404, { error: 'not found' });
} catch (e) { send(500, { error: e.message }); }
});
server.listen(PORT, '127.0.0.1', () => console.log(`[bridge] in ascolto su 127.0.0.1:${PORT}`));
start().catch((e) => { state.error = e.message; console.error(e); });
+1603
View File
File diff suppressed because it is too large. Load diff
+12
View File
@@ -0,0 +1,12 @@
{
"name": "avatarpy-whatsapp-bridge",
"version": "1.0.0",
"private": true,
"type": "module",
"main": "index.js",
"dependencies": {
"@whiskeysockets/baileys": "^6.7.0",
"@hapi/boom": "^10.0.1",
"pino": "^9.0.0"
}
}