Compare commits

..
7 Commits
Author SHA1 Message Date
luciano 39a117f08b feat: cookies-from-browser per bypass HTTP 403 YouTube (v1.10.1)
Alcuni video YouTube restituiscono 403 anche con yt-dlp aggiornato.
Fix: nuova opzione Impostazioni "Cookies da browser" con dropdown
(chrome/safari/firefox/edge/brave). Priorita': file cookies.txt
esistente -> --cookies-from-browser <name> -> nessuno.

Bonus:
- core/paths.py::_bundle_dirs() ora include <project>/bundle_bin/
  anche in dev, cosi' python main.py usa lo stesso yt-dlp del build
  invece di Homebrew (spesso obsoleto)
- Helper _cookie_args() in downloader.py per non ripetere la logica
  su tutti i punti dove serve --cookies
2026-08-19 15:30:46 +02:00
luciano e7d11d0235 feat: Dedup enhancements (streaming, retry, filename mode, scroll, auto-select)
Iterazioni sul feature Dedup (v1.10.0) da subito-post-scaffolding:
- compute_fingerprint espone il vero errore invece del generico None,
  worker propaga l'err_msg via progress_callback esteso
- Retry automatico: -length 30 su "invalid data / decoding frame",
  -length 60 + -algorithm 1 su "fingerprint vuoto"
- Nuovo metodo "filename" (Jaccard sui token nome file) con algoritmo
  incrementale O(N·K) — utile per file corrotti o per anteprima veloce
- Streaming groups: sia fingerprint che filename emettono `dedup:group`
  appena un gruppo raggiunge >=2 file; JS accumula in Map per update
  in-place. Progress bar avanza in tempo reale
- Fix shape bug che dava "NaN duplicati" nella summary line
- .dedup-groups-scroll: max-height 55vh + overflow-y auto per non
  perdere l'header/footer scrollando molti gruppi
- Bottone "Seleziona tutti i consigliati" nel footer: ripristina la
  selezione di default (tutti tranne il TIENI marcati per cancellazione)
- .gitignore: dedup_cache.db (SQLite locale per macchina utente)
2026-08-02 13:59:39 +02:00
luciano 00ff5a6220 ui: DedupUI + stili gruppi duplicati
- DedupUI: state locale + folder picker + start/stop scan +
  render gruppi (checkbox multi-select), file col bitrate max
  auto-preselezionato come "TIENI" (verde, checkbox disabled),
  audio preview riusa _makePreviewBtn dal picker Upgrade.
- Listener bridge dedup:progress / dedup:done, ripristino
  ultima cartella + recursive da config al mount.
- CSS: .dedup-group-card, .dedup-file-row, .dedup-keep (verde),
  .dedup-keep-badge, .dedup-footer-bar sticky.
2026-08-02 13:23:04 +02:00
luciano e90e38edc8 ui: markup tab Dedup 2026-08-02 13:22:52 +02:00
luciano f18dd375e8 bridge: metodi Api per tab Dedup + bump v1.10.0
- Api.dedup_pick_folder / dedup_start_scan / dedup_stop_scan
  / dedup_move_to_trash con worker thread + eventi
  dedup:progress e dedup:done.
- core.config: dedup_last_folder, dedup_recursive defaults;
  VERSION -> v1.10.0 (minor: nuova tab Dedup).
2026-08-02 13:22:40 +02:00
luciano fb13b0f70e dedup: modulo audio fingerprinting via fpcalc + cache SQLite
Nuovo modulo core.dedup per rilevare brani audio duplicati usando
Chromaprint (fpcalc): scan cartella (opzionalmente ricorsivo),
fingerprint acustico via fpcalc -json, cache SQLite in
_get_config_dir()/dedup_cache.db per non ricalcolare al re-scan,
raggruppamento per fingerprint identico + ordering per bitrate DESC.
move_to_trash() usa send2trash (reversibile via Finder/Explorer).

- core/paths.py: find_fpcalc() con fallback dev su project bundle_bin
- core/dedup.py: scan_folder, compute_fingerprint, move_to_trash
- tests/test_dedup.py: 9 unit test (mock fpcalc + send2trash)
- requirements.txt: send2trash>=1.8.0
- build_macos.py: download fpcalc universal binary da GitHub releases
- build_windows.py: download fpcalc.exe da GitHub releases
2026-08-02 13:22:24 +02:00
luciano e09ee1382b fix: typo _find_ffmpeg_dir in update_cover_only (bump v1.9.5)
Il typo (underscore prefix) causava NameError nel thread upgrade
quando un file era già HQ (soglia superata) e serviva solo aggiornare
la copertina. Il thread moriva silenziosamente e l'utente vedeva
l'app 'bloccata' senza feedback.
2026-08-02 13:02:26 +02:00
14 changed files with 1591 additions and 16 deletions

No files matched your search

+4
View File
@@ -49,3 +49,7 @@ server/node_modules/
server/.wrangler/ server/.wrangler/
server/.dev.vars server/.dev.vars
server/dist/ server/dist/
# Dedup local cache (SQLite fingerprint cache, per macchina dell'utente)
dedup_cache.db
dedup_cache.db-journal
+143
View File
@@ -46,6 +46,7 @@ from core.upgrader import (
request_stop as request_upgrade_stop, request_stop as request_upgrade_stop,
count_files_info, count_files_info,
) )
from core import dedup
SPOTIFY_GUIDE_TEXT = """\ SPOTIFY_GUIDE_TEXT = """\
@@ -111,6 +112,7 @@ class Api:
self._download_thread: Optional[threading.Thread] = None self._download_thread: Optional[threading.Thread] = None
self._upgrade_thread: Optional[threading.Thread] = None self._upgrade_thread: Optional[threading.Thread] = None
self._video_thread: Optional[threading.Thread] = None self._video_thread: Optional[threading.Thread] = None
self._dedup_thread: Optional[threading.Thread] = None
# Coda usata dal resolve_callback per attendere la scelta utente # Coda usata dal resolve_callback per attendere la scelta utente
# sul modal "match locali multipli" della tab Upgrade. # sul modal "match locali multipli" della tab Upgrade.
self._upgrade_resolve_q: "queue.Queue[dict]" = queue.Queue(1) self._upgrade_resolve_q: "queue.Queue[dict]" = queue.Queue(1)
@@ -168,6 +170,7 @@ class Api:
"bitrate": payload.get("bitrate", "320K"), "bitrate": payload.get("bitrate", "320K"),
"hq_threshold": threshold, "hq_threshold": threshold,
"cookies_path": (payload.get("cookies_path") or "").strip(), "cookies_path": (payload.get("cookies_path") or "").strip(),
"cookies_browser": (payload.get("cookies_browser") or "").strip().lower(),
"output_dir": (payload.get("output_dir") or "").strip(), "output_dir": (payload.get("output_dir") or "").strip(),
"theme": payload.get("theme", "dark"), "theme": payload.get("theme", "dark"),
}) })
@@ -1820,3 +1823,143 @@ class Api:
}) })
self._emit("convert:done", {"ok": True}) self._emit("convert:done", {"ok": True})
# ================================================================
# DEDUP — audio duplicati via Chromaprint fingerprinting
# ================================================================
def dedup_pick_folder(self) -> str:
"""Folder picker per la cartella da scansionare."""
return self.browse_directory()
def dedup_start_scan(self, payload: dict) -> dict:
"""Avvia worker di scansione. Payload: {directory, recursive}.
Salva `dedup_last_folder` e `dedup_recursive` in config per la
prossima apertura del tab.
"""
if self._dedup_thread and self._dedup_thread.is_alive():
return {"ok": False, "error": "Scansione dedup gia in corso"}
directory = (payload.get("directory") or "").strip()
recursive = bool(payload.get("recursive", True))
method = (payload.get("method") or "fingerprint").strip()
if method not in ("fingerprint", "filename"):
method = "fingerprint"
if not directory:
return {"ok": False, "error": "Cartella non impostata"}
if not os.path.isdir(directory):
return {"ok": False, "error": "Cartella non trovata"}
# Persist last folder/recursive/method
try:
cfg = load_config()
cfg["dedup_last_folder"] = directory
cfg["dedup_recursive"] = recursive
cfg["dedup_method"] = method
save_config(cfg)
except Exception:
pass
self._dedup_thread = threading.Thread(
target=self._dedup_worker,
args=(directory, recursive, method),
daemon=True,
)
self._dedup_thread.start()
return {"ok": True}
def dedup_stop_scan(self) -> dict:
dedup.request_stop()
self._log("dedup", "[INFO] Interruzione richiesta...")
return {"ok": True}
def dedup_move_to_trash(self, paths: list) -> dict:
"""Sposta i file in cestino via send2trash. Non consuma quota."""
if not isinstance(paths, list):
return {"ok": False, "error": "paths deve essere una lista"}
# Sanitize: solo str non vuote
clean = [str(p).strip() for p in paths if p and str(p).strip()]
if not clean:
return {"ok": False, "error": "Nessun file da cancellare"}
result = dedup.move_to_trash(clean)
moved_n = len(result.get("moved", []))
failed_n = len(result.get("failed", []))
if moved_n:
self._log("dedup", f"[OK] {moved_n} file spostati nel cestino")
for f in result.get("failed", []):
self._log("dedup",
f"[ERRORE] {f.get('path')}: {f.get('error')}")
return {"ok": True, "moved": result.get("moved", []),
"failed": result.get("failed", []),
"moved_count": moved_n, "failed_count": failed_n}
def _dedup_worker(self, directory: str, recursive: bool,
method: str = "fingerprint") -> None:
"""Esegue la scansione in background e emette progress/done."""
view = "dedup"
dedup.reset_stop()
method_label = "audio fingerprint" if method == "fingerprint" else "nome file"
self._log(view, f"[INFO] Scansione: {directory} (metodo: {method_label}, recursive={recursive})")
_last = [0.0]
_THROTTLE = 0.05
def progress_cb(idx: int, total: int, filename: str, status: str, err_msg: str = "") -> None:
# Throttle solo eventi 'computing'/'cached' (che possono essere migliaia)
if status in ("computing", "cached"):
now = time.monotonic()
if now - _last[0] < _THROTTLE and idx != total:
return
_last[0] = now
payload_evt = {
"idx": idx, "total": total,
"filename": filename, "status": status,
}
if err_msg:
payload_evt["error_msg"] = err_msg
if total > 0:
payload_evt["overall"] = min(idx / total, 1.0)
if status == "completed":
payload_evt["overall"] = 1.0
self._log(view, f"[INFO] Scansione completata ({total} file).")
elif status == "stopped":
self._log(view, "[INFO] Scansione interrotta.")
elif status == "error" and filename:
detail = f": {err_msg}" if err_msg else ""
self._log(view, f"[ERRORE] {filename}{detail}")
self._emit("dedup:progress", payload_evt)
def group_cb(group: dict) -> None:
"""Emit streaming: appena un gruppo raggiunge/aggiorna >=2 file."""
try:
self._emit("dedup:group", group)
except Exception:
pass
try:
groups = dedup.scan_folder(directory, recursive=recursive,
progress_callback=progress_cb,
method=method,
group_callback=group_cb)
except Exception as e:
self._log(view, f"[ERRORE] {e}")
self._emit("dedup:done", {"ok": False, "error": str(e),
"groups": []})
return
n_groups = len(groups)
n_dupes = sum(max(0, len(g) - 1) for g in groups)
total_bytes = sum(sum(int(e.get("size") or 0) for e in g[1:])
for g in groups)
self._log(view,
f"[INFO] Gruppi: {n_groups} — duplicati: {n_dupes} — "
f"spazio recuperabile: ~{total_bytes // 1024 // 1024} MB")
self._emit("dedup:done", {
"ok": True,
"groups": groups,
"n_groups": n_groups,
"n_dupes": n_dupes,
"reclaimable_bytes": total_bytes,
})
+43
View File
@@ -64,6 +64,48 @@ def download_ytdlp():
return dest return dest
# =========================================================================
# 1b. Scarica fpcalc (Chromaprint) — usato dal tab Dedup
# =========================================================================
def download_fpcalc():
"""Scarica il binario universal Chromaprint fpcalc per macOS."""
dest = BUNDLE_DIR / "fpcalc"
if dest.exists():
log(f"fpcalc gia presente: {dest}")
return dest
url = ("https://github.com/acoustid/chromaprint/releases/download/"
"v1.5.1/chromaprint-fpcalc-1.5.1-macos-universal.tar.gz")
log(f"Scarico fpcalc da {url} ...")
BUNDLE_DIR.mkdir(parents=True, exist_ok=True)
tmp_tar = BUNDLE_DIR / "_fpcalc.tar.gz"
subprocess.run(["curl", "-L", "-o", str(tmp_tar), url], check=True)
# Estrae ovunque nella cartella, poi trova fpcalc e lo sposta al posto
tmp_extract = BUNDLE_DIR / "_fpcalc_extract"
if tmp_extract.exists():
shutil.rmtree(tmp_extract)
tmp_extract.mkdir()
subprocess.run(["tar", "-xzf", str(tmp_tar), "-C", str(tmp_extract)],
check=True)
# Trova fpcalc nel folder estratto
found = None
for p in tmp_extract.rglob("fpcalc"):
if p.is_file():
found = p
break
if not found:
shutil.rmtree(tmp_extract, ignore_errors=True)
tmp_tar.unlink(missing_ok=True)
raise RuntimeError("fpcalc non trovato nell'archivio")
shutil.copy2(found, dest)
dest.chmod(dest.stat().st_mode | stat.S_IEXEC)
shutil.rmtree(tmp_extract, ignore_errors=True)
tmp_tar.unlink(missing_ok=True)
log(f"fpcalc scaricato: {dest}")
return dest
# ========================================================================= # =========================================================================
# 2. Raccogli ffmpeg/ffprobe + dylib # 2. Raccogli ffmpeg/ffprobe + dylib
# ========================================================================= # =========================================================================
@@ -511,6 +553,7 @@ def main():
shutil.rmtree(BUNDLE_DIR) shutil.rmtree(BUNDLE_DIR)
download_ytdlp() download_ytdlp()
download_fpcalc()
bundle_ffmpeg() bundle_ffmpeg()
run_pyinstaller() run_pyinstaller()
+30
View File
@@ -31,6 +31,8 @@ BUNDLE_DIR = ROOT / "bundle_bin"
YTDLP_URL = "https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp.exe" YTDLP_URL = "https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp.exe"
FFMPEG_URL = "https://github.com/BtbN/FFmpeg-Builds/releases/download/latest/ffmpeg-master-latest-win64-gpl.zip" FFMPEG_URL = "https://github.com/BtbN/FFmpeg-Builds/releases/download/latest/ffmpeg-master-latest-win64-gpl.zip"
FPCALC_URL = ("https://github.com/acoustid/chromaprint/releases/download/"
"v1.5.1/chromaprint-fpcalc-1.5.1-windows-x86_64.zip")
def log(msg): def log(msg):
@@ -87,6 +89,33 @@ def download_ffmpeg():
sys.exit(1) sys.exit(1)
# =========================================================================
# 2b. Scarica fpcalc.exe (Chromaprint) — usato dal tab Dedup
# =========================================================================
def download_fpcalc():
dest = BUNDLE_DIR / "fpcalc.exe"
if dest.exists():
log("fpcalc.exe gia presente")
return
log(f"Scarico fpcalc.exe ...")
BUNDLE_DIR.mkdir(parents=True, exist_ok=True)
response = urllib.request.urlopen(FPCALC_URL)
zip_data = io.BytesIO(response.read())
with zipfile.ZipFile(zip_data) as zf:
for member in zf.namelist():
basename = Path(member).name
if basename == "fpcalc.exe":
data = zf.read(member)
dest.write_bytes(data)
log(f" Estratto: fpcalc.exe ({len(data) // 1024} KB)")
if not dest.exists():
print("ERRORE: fpcalc.exe non trovato nello zip!")
sys.exit(1)
# ========================================================================= # =========================================================================
# 3. PyInstaller # 3. PyInstaller
# ========================================================================= # =========================================================================
@@ -187,6 +216,7 @@ def main():
download_ytdlp() download_ytdlp()
download_ffmpeg() download_ffmpeg()
download_fpcalc()
run_pyinstaller() run_pyinstaller()
print("\n" + "=" * 50) print("\n" + "=" * 50)
+6 -1
View File
@@ -5,7 +5,7 @@ import os
import sys import sys
from pathlib import Path from pathlib import Path
VERSION = "v1.9.4" VERSION = "v1.10.1"
APP_NAME = "MusicTools" APP_NAME = "MusicTools"
@@ -68,6 +68,7 @@ DEFAULTS = {
"bitrate": "320K", "bitrate": "320K",
"hq_threshold": 310, "hq_threshold": 310,
"cookies_path": str(_project_dir / "cookies.txt"), "cookies_path": str(_project_dir / "cookies.txt"),
"cookies_browser": "", # "" | chrome | safari | firefox | edge | brave
"output_dir": str(_project_dir / "MUSICA"), "output_dir": str(_project_dir / "MUSICA"),
"theme": "dark", "theme": "dark",
# ---- Beatport ---- # ---- Beatport ----
@@ -78,6 +79,10 @@ DEFAULTS = {
"spotify_search_last_query": "", "spotify_search_last_query": "",
"spotify_search_artist_mode": False, "spotify_search_artist_mode": False,
"youtube_search_last_query": "", "youtube_search_last_query": "",
# ---- Dedup (audio duplicati via Chromaprint) ----
"dedup_last_folder": "",
"dedup_recursive": True,
"dedup_method": "fingerprint", # "fingerprint" | "filename"
# ---- Licenza ---- # ---- Licenza ----
"license_key": "", # chiave fornita all'utente via email "license_key": "", # chiave fornita all'utente via email
"license_email": "", # email associata all'acquisto "license_email": "", # email associata all'acquisto
+518
View File
@@ -0,0 +1,518 @@
"""Deduplicator audio via Chromaprint fingerprinting + SQLite cache.
Pipeline:
1. Scansiona la cartella (opzionalmente ricorsivo) filtrando per
estensioni audio (AUDIO_EXTENSIONS di core.upgrader).
2. Per ogni file calcola il fingerprint Chromaprint (`fpcalc -json`).
Il valore viene messo in cache SQLite: al re-scan, se
(size, mtime) coincide col record, riusiamo il fingerprint senza
rilanciare fpcalc.
3. Raggruppa i file per fingerprint identico (>= 2 file). Per ogni
gruppo, i file vengono ordinati per bitrate DESC (tie-break: size
DESC): il primo e' quello "da tenere", gli altri i duplicati.
4. `move_to_trash` invia i path selezionati al cestino di sistema
tramite send2trash (reversibile via Finder/Explorer).
Progress callback firma:
(processed: int, total: int, filename: str, status: str)
Status validi: 'scanning' | 'computing' | 'cached' | 'error' | 'stopped'
| 'completed'.
"""
from __future__ import annotations
import json
import sqlite3
import subprocess
import threading
from pathlib import Path
from typing import Callable, Optional
from core.paths import find_fpcalc, subprocess_flags
from core.upgrader import AUDIO_EXTENSIONS, get_bitrate
# Timeout massimo per una singola invocazione fpcalc.
_FPCALC_TIMEOUT_SEC = 30
# ------------------------------------------------------------------
# Stop / interrupt
# ------------------------------------------------------------------
_stop_event = threading.Event()
def request_stop() -> None:
"""Segnala al worker di interrompere la scansione al prossimo file."""
_stop_event.set()
def reset_stop() -> None:
"""Azzera il flag di stop prima di iniziare una nuova scansione."""
_stop_event.clear()
def is_stopped() -> bool:
return _stop_event.is_set()
# ------------------------------------------------------------------
# Cache SQLite
# ------------------------------------------------------------------
def _cache_db_path() -> Path:
"""Path del DB di cache dei fingerprint.
Riusa `_get_config_dir` di core.config cosi' finisce nella stessa
cartella di config.json (~/Library/Application Support/MusicTools/
su macOS, %APPDATA%/MusicTools/ su Windows, project root in dev).
"""
from core.config import _get_config_dir
return _get_config_dir() / "dedup_cache.db"
def _init_db(conn: sqlite3.Connection) -> None:
"""Crea (idempotente) lo schema della cache."""
conn.execute(
"""
CREATE TABLE IF NOT EXISTS files (
path TEXT PRIMARY KEY,
size INTEGER NOT NULL,
mtime REAL NOT NULL,
duration REAL,
fingerprint TEXT,
bitrate INTEGER
)
"""
)
conn.commit()
def _open_cache(db_path: Optional[Path] = None) -> sqlite3.Connection:
"""Apre (creando se serve) la connessione alla cache."""
p = db_path or _cache_db_path()
p.parent.mkdir(parents=True, exist_ok=True)
conn = sqlite3.connect(str(p))
_init_db(conn)
return conn
def _cache_get(conn: sqlite3.Connection, path: str,
size: int, mtime: float) -> Optional[dict]:
"""Ritorna il record se (size, mtime) invariato, altrimenti None."""
cur = conn.execute(
"SELECT size, mtime, duration, fingerprint, bitrate FROM files WHERE path = ?",
(path,),
)
row = cur.fetchone()
if not row:
return None
csize, cmtime, dur, fp, br = row
# Tolleranza minima sul mtime (float precision su alcuni FS)
if csize != size or abs(float(cmtime) - float(mtime)) > 0.001:
return None
if not fp:
return None
return {
"size": int(csize),
"mtime": float(cmtime),
"duration": float(dur) if dur is not None else 0.0,
"fingerprint": str(fp),
"bitrate": int(br) if br is not None else 0,
}
def _cache_put(conn: sqlite3.Connection, path: str, size: int, mtime: float,
duration: float, fingerprint: str, bitrate: int) -> None:
"""Upsert (SQLite ha ON CONFLICT REPLACE via INSERT OR REPLACE)."""
conn.execute(
"INSERT OR REPLACE INTO files (path, size, mtime, duration, fingerprint, bitrate)"
" VALUES (?, ?, ?, ?, ?, ?)",
(path, int(size), float(mtime), float(duration or 0),
str(fingerprint or ""), int(bitrate or 0)),
)
conn.commit()
# ------------------------------------------------------------------
# fpcalc
# ------------------------------------------------------------------
def _run_fpcalc(fpcalc: str, path: str, length: Optional[int] = None) -> dict:
"""Esegue fpcalc una volta. Ritorna {duration, fingerprint} su successo
o {_error: str} su fallimento."""
cmd = [fpcalc, "-json"]
if length is not None:
cmd += ["-length", str(length)]
cmd.append(str(path))
try:
proc = subprocess.run(
cmd,
capture_output=True,
text=True,
timeout=_FPCALC_TIMEOUT_SEC,
**subprocess_flags(),
)
except subprocess.TimeoutExpired:
return {"_error": f"timeout {_FPCALC_TIMEOUT_SEC}s"}
except (OSError, ValueError) as e:
return {"_error": f"subprocess: {e}"}
if proc.returncode != 0:
err = (proc.stderr or proc.stdout or "").strip().splitlines()
msg = err[-1] if err else f"exit {proc.returncode}"
return {"_error": msg[:200]}
try:
data = json.loads(proc.stdout or "{}")
except (json.JSONDecodeError, ValueError) as e:
return {"_error": f"JSON malformato: {e}"}
fp = data.get("fingerprint")
if not fp:
return {"_error": "fingerprint vuoto (audio troppo corto?)"}
try:
dur = float(data.get("duration") or 0)
except (TypeError, ValueError):
dur = 0.0
return {"duration": dur, "fingerprint": str(fp)}
def compute_fingerprint(fpcalc: str, path: str) -> Optional[dict]:
"""Chiama fpcalc e ritorna {duration, fingerprint} o {_error}.
Se il primo tentativo (full length) fallisce con "Invalid data" o simili
(frame audio corrotti che libav rifiuta), riprova con `-length 30`.
Molti file danneggiati hanno i frame corrotti nella parte finale e
limitando la scansione ai primi 30s si riesce a estrarre comunque
un fingerprint affidabile (30s bastano per l'unicità Chromaprint).
"""
if not fpcalc:
return {"_error": "fpcalc non trovato nel bundle"}
res = _run_fpcalc(fpcalc, path)
if "fingerprint" in res:
return res
err_msg = res.get("_error", "").lower()
# Retry 1: frame audio corrotti → riduci finestra a 30s
corrupt_signals = ("invalid data", "decoding audio frame",
"error while decoding", "invalid frame")
if any(sig in err_msg for sig in corrupt_signals):
res2 = _run_fpcalc(fpcalc, path, length=30)
if "fingerprint" in res2:
res2["_partial"] = True # 30s soltanto
return res2
# Retry 2: fingerprint vuoto → prova con finestra piu' lunga (60s)
# nel caso l'intro sia silenzio/muto (chromaprint richiede audio "reale")
if "vuoto" in err_msg or "empty" in err_msg:
res2 = _run_fpcalc(fpcalc, path, length=60)
if "fingerprint" in res2:
res2["_partial"] = True
return res2
# Ancora vuoto → prova algoritmo differente (chromaprint algo 1)
# tramite subprocess diretto perche' _run_fpcalc non lo supporta
try:
proc = subprocess.run(
[fpcalc, "-json", "-length", "60", "-algorithm", "1", str(path)],
capture_output=True, text=True,
timeout=_FPCALC_TIMEOUT_SEC,
**subprocess_flags(),
)
if proc.returncode == 0:
data = json.loads(proc.stdout or "{}")
fp = data.get("fingerprint")
if fp:
return {
"duration": float(data.get("duration") or 0),
"fingerprint": str(fp),
"_partial": True,
}
except Exception:
pass
return res
# ------------------------------------------------------------------
# Scan
# ------------------------------------------------------------------
def _iter_audio_files(directory: str, recursive: bool) -> list[Path]:
"""Elenca tutti i file audio (estensione case-insensitive)."""
base = Path(directory)
if not base.exists() or not base.is_dir():
return []
files: list[Path] = []
if recursive:
for f in base.rglob("*"):
if f.is_file() and f.suffix.lower() in AUDIO_EXTENSIONS:
files.append(f)
else:
for f in base.iterdir():
if f.is_file() and f.suffix.lower() in AUDIO_EXTENSIONS:
files.append(f)
files.sort()
return files
def _scan_by_filename(files: list, progress_callback: Optional[Callable],
group_callback: Optional[Callable] = None,
similarity_threshold: float = 0.8) -> list[list[dict]]:
"""Raggruppa file per similarità nome (Jaccard sui token normalizzati),
algoritmo INCREMENTALE: per ogni nuovo file cerca match tra i gruppi già
formati (lookup O(K) dove K = numero gruppi). Emette streaming via
`group_callback` appena un gruppo raggiunge ≥ 2 file.
"""
from core.upgrader import _normalize_stem # riuso
def _pc(idx, total_n, name, status, err=""):
if not progress_callback:
return
try:
progress_callback(idx, total_n, name, status, err)
except TypeError:
progress_callback(idx, total_n, name, status)
def _gc(group_id: str, entries: list) -> None:
if group_callback:
try:
group_callback({"id": group_id, "entries": list(entries)})
except Exception:
pass
total = len(files)
# Ogni voce: {"id": str, "key_tokens": frozenset, "entries": [dict]}
groups: list = []
for i, fp_path in enumerate(files, start=1):
if is_stopped():
_pc(i - 1, total, "", "stopped")
break
try:
size = fp_path.stat().st_size
except OSError as e:
_pc(i, total, fp_path.name, "error", f"stat: {e}")
continue
tokens = frozenset(_normalize_stem(fp_path.stem))
if not tokens:
_pc(i, total, fp_path.name, "error", "nome senza token utili")
continue
try:
bitrate = get_bitrate(fp_path)
except Exception:
bitrate = 0
entry = {
"path": str(fp_path), "size": size, "bitrate": bitrate,
"duration": 0, "fingerprint": "",
}
# Cerca match nei gruppi già formati (lineare sui gruppi, non sui file)
matched = None
for g in groups:
common = len(tokens & g["key_tokens"])
if common == 0:
continue
union = len(tokens | g["key_tokens"])
if union > 0 and (common / union) >= similarity_threshold:
matched = g
break
if matched is not None:
was_solo = len(matched["entries"]) == 1
matched["entries"].append(entry)
matched["entries"].sort(key=lambda e: (-e["bitrate"], -e["size"]))
# Streaming: emit ogni volta che il gruppo diventa/rimane ≥ 2 file
_gc(matched["id"], matched["entries"])
else:
gid = f"fn_{len(groups)}_{fp_path.stem[:20]}"
groups.append({
"id": gid,
"key_tokens": set(tokens),
"entries": [entry],
})
_pc(i, total, fp_path.name, "cached")
# Ritorna solo i gruppi con >= 2 file
result = [sorted(g["entries"], key=lambda e: (-e["bitrate"], -e["size"]))
for g in groups if len(g["entries"]) >= 2]
result.sort(key=lambda g: -max(e["size"] for e in g))
_pc(total, total, "", "completed")
return result
def scan_folder(
directory: str,
recursive: bool = True,
progress_callback: Optional[Callable] = None,
method: str = "fingerprint",
group_callback: Optional[Callable] = None,
) -> list[list[dict]]:
"""Ritorna la lista di gruppi di file duplicati (>= 2 file).
Ogni file nel gruppo e' un dict:
{path, size, bitrate, duration, fingerprint}
Gruppi ordinati per size del file piu' grande DESC (i gruppi che
occupano piu' spazio vengono prima). All'interno di ogni gruppo:
bitrate DESC, poi size DESC (il primo e' quello "da tenere").
`method`:
- "fingerprint" (default): Chromaprint via fpcalc, preciso ma lento.
Cache SQLite persistente. Raggruppa per fingerprint identico.
- "filename": similarità Jaccard sui nomi file. Veloce ma euristico.
Non richiede fpcalc.
"""
reset_stop()
files = _iter_audio_files(directory, recursive)
total = len(files)
if total == 0:
if progress_callback:
progress_callback(0, 0, "", "completed")
return []
if method == "filename":
return _scan_by_filename(files, progress_callback, group_callback)
fpcalc = find_fpcalc()
if not fpcalc:
# Senza fpcalc non possiamo fare nulla. Segnaliamo errore su ogni
# file e ritorniamo lista vuota.
if progress_callback:
progress_callback(0, total, "", "error")
return []
conn = _open_cache()
try:
# {fingerprint: [entry, ...]}
by_fp: dict[str, list[dict]] = {}
def _pc(idx, total_n, name, status, err=""):
"""Chiama progress_callback in modo retrocompatibile: la firma
legacy è a 4 args, quella nuova a 5 con `error_msg` opzionale."""
if not progress_callback:
return
try:
progress_callback(idx, total_n, name, status, err)
except TypeError:
progress_callback(idx, total_n, name, status)
for i, fp_path in enumerate(files, start=1):
if is_stopped():
_pc(i - 1, total, "", "stopped")
return []
try:
st = fp_path.stat()
size = st.st_size
mtime = st.st_mtime
except OSError as e:
_pc(i, total, fp_path.name, "error", f"stat: {e}")
continue
path_str = str(fp_path)
cached = _cache_get(conn, path_str, size, mtime)
if cached:
fp_hash = cached["fingerprint"]
duration = cached["duration"]
bitrate = cached["bitrate"] or get_bitrate(fp_path)
_pc(i, total, fp_path.name, "cached")
else:
_pc(i, total, fp_path.name, "computing")
res = compute_fingerprint(fpcalc, path_str)
if not res or not res.get("fingerprint"):
err_msg = (res or {}).get("_error", "errore sconosciuto")
_pc(i, total, fp_path.name, "error", err_msg)
continue
fp_hash = res["fingerprint"]
duration = res["duration"]
try:
bitrate = get_bitrate(fp_path)
except Exception:
bitrate = 0
_cache_put(conn, path_str, size, mtime, duration, fp_hash, bitrate)
entry = {
"path": path_str,
"size": int(size),
"bitrate": int(bitrate or 0),
"duration": float(duration or 0),
"fingerprint": fp_hash,
}
grp = by_fp.setdefault(fp_hash, [])
grp.append(entry)
# Streaming: appena il gruppo raggiunge (o supera) 2 elementi,
# emetti update (JS accumula/aggiorna in tempo reale)
if group_callback and len(grp) >= 2:
# Ordinamento intra-gruppo prima di emit (best-to-keep primo)
grp.sort(key=lambda e: (-int(e.get("bitrate") or 0),
-int(e.get("size") or 0)))
try:
group_callback({
"id": f"fp_{fp_hash[:24]}",
"entries": list(grp),
})
except Exception:
pass
finally:
try:
conn.close()
except Exception:
pass
# Filtra: solo gruppi con >= 2 file
groups = [g for g in by_fp.values() if len(g) >= 2]
# Sort dei file dentro il gruppo: bitrate DESC, size DESC.
# Sort dei gruppi: size del file piu' grande DESC (usa max del gruppo).
for g in groups:
g.sort(key=lambda e: (-int(e.get("bitrate") or 0),
-int(e.get("size") or 0)))
groups.sort(key=lambda g: -max(int(e.get("size") or 0) for e in g))
if progress_callback:
progress_callback(total, total, "", "completed")
return groups
# ------------------------------------------------------------------
# Trash
# ------------------------------------------------------------------
def move_to_trash(paths: list[str]) -> dict:
"""Sposta i file in cestino tramite send2trash.
Ritorna {moved: [...], failed: [{path, error}, ...]}. Non solleva
mai eccezioni: gli errori per file singolo finiscono in `failed`.
Aggiorna la cache SQLite rimuovendo i record dei file spostati (per
quelli riusciti), cosi' un re-scan non li propone piu'.
"""
# Import interno per rendere il modulo importabile anche se
# send2trash non e' installato (i test possono mockarlo).
try:
from send2trash import send2trash
except Exception as e: # pragma: no cover — solo se pacchetto mancante
return {
"moved": [],
"failed": [{"path": p, "error": f"send2trash non disponibile: {e}"}
for p in (paths or [])],
}
moved: list[str] = []
failed: list[dict] = []
for p in (paths or []):
try:
send2trash(p)
moved.append(p)
except Exception as e:
failed.append({"path": p, "error": str(e)})
# Cache cleanup best-effort (non fatale se fallisce)
if moved:
try:
conn = _open_cache()
try:
for p in moved:
conn.execute("DELETE FROM files WHERE path = ?", (p,))
conn.commit()
finally:
conn.close()
except Exception:
pass
return {"moved": moved, "failed": failed}
+30 -14
View File
@@ -12,6 +12,29 @@ from typing import Callable, Optional
from core.paths import find_ytdlp, find_ffmpeg_dir, subprocess_flags from core.paths import find_ytdlp, find_ffmpeg_dir, subprocess_flags
_ALLOWED_BROWSERS = {"chrome", "safari", "firefox", "edge", "brave", "chromium", "opera", "vivaldi"}
def _cookie_args(cookies_path: Optional[str]) -> list:
"""Ritorna gli argomenti yt-dlp per i cookies.
Priorità: file cookies_path se esiste → altrimenti --cookies-from-browser
<name> se `cookies_browser` è settato in config → altrimenti niente.
"""
if cookies_path and Path(cookies_path).exists():
return ["--cookies", cookies_path]
# Fallback: legge il browser dalla config al volo (evita di cambiare
# firma di tutte le funzioni download_*)
try:
from core.config import load_config
browser = (load_config().get("cookies_browser") or "").strip().lower()
except Exception:
browser = ""
if browser in _ALLOWED_BROWSERS:
return ["--cookies-from-browser", browser]
return []
# Flag globale per interruzione # Flag globale per interruzione
_stop_event = threading.Event() _stop_event = threading.Event()
_current_process: Optional[subprocess.Popen] = None _current_process: Optional[subprocess.Popen] = None
@@ -99,8 +122,7 @@ def _search_youtube(query: str, cookies_path: Optional[str] = None) -> tuple[str
"--no-warnings", "--no-warnings",
"--flat-playlist", "--flat-playlist",
] ]
if cookies_path and Path(cookies_path).exists(): cmd.extend(_cookie_args(cookies_path))
cmd.extend(["--cookies", cookies_path])
result = subprocess.run(cmd, capture_output=True, text=True, timeout=30, **subprocess_flags()) result = subprocess.run(cmd, capture_output=True, text=True, timeout=30, **subprocess_flags())
if result.returncode != 0: if result.returncode != 0:
@@ -216,8 +238,7 @@ def download_playlist(
ffmpeg_dir = find_ffmpeg_dir() ffmpeg_dir = find_ffmpeg_dir()
if ffmpeg_dir: if ffmpeg_dir:
cmd.extend(["--ffmpeg-location", ffmpeg_dir]) cmd.extend(["--ffmpeg-location", ffmpeg_dir])
if cookies_path and Path(cookies_path).exists(): cmd.extend(_cookie_args(cookies_path))
cmd.extend(["--cookies", cookies_path])
try: try:
with _process_lock: with _process_lock:
@@ -300,8 +321,7 @@ def download_direct_url(
"--no-warnings", "--no-warnings",
url, url,
] ]
if cookies_path and Path(cookies_path).exists(): probe_cmd.extend(_cookie_args(cookies_path))
probe_cmd.extend(["--cookies", cookies_path])
try: try:
result = subprocess.run(probe_cmd, capture_output=True, text=True, timeout=60, **subprocess_flags()) result = subprocess.run(probe_cmd, capture_output=True, text=True, timeout=60, **subprocess_flags())
@@ -375,8 +395,7 @@ def download_direct_url(
ffmpeg_dir = find_ffmpeg_dir() ffmpeg_dir = find_ffmpeg_dir()
if ffmpeg_dir: if ffmpeg_dir:
cmd.extend(["--ffmpeg-location", ffmpeg_dir]) cmd.extend(["--ffmpeg-location", ffmpeg_dir])
if cookies_path and Path(cookies_path).exists(): cmd.extend(_cookie_args(cookies_path))
cmd.extend(["--cookies", cookies_path])
try: try:
with _process_lock: with _process_lock:
@@ -511,8 +530,7 @@ def download_urls(
ffmpeg_dir = find_ffmpeg_dir() ffmpeg_dir = find_ffmpeg_dir()
if ffmpeg_dir: if ffmpeg_dir:
cmd.extend(["--ffmpeg-location", ffmpeg_dir]) cmd.extend(["--ffmpeg-location", ffmpeg_dir])
if cookies_path and Path(cookies_path).exists(): cmd.extend(_cookie_args(cookies_path))
cmd.extend(["--cookies", cookies_path])
try: try:
with _process_lock: with _process_lock:
@@ -608,8 +626,7 @@ def download_video(
probe_cmd = [ probe_cmd = [
ytdlp, "--dump-json", "--flat-playlist", "--no-download", "--no-warnings", url, ytdlp, "--dump-json", "--flat-playlist", "--no-download", "--no-warnings", url,
] ]
if cookies_path and Path(cookies_path).exists(): probe_cmd.extend(_cookie_args(cookies_path))
probe_cmd.extend(["--cookies", cookies_path])
try: try:
result = subprocess.run(probe_cmd, capture_output=True, text=True, timeout=60, **subprocess_flags()) result = subprocess.run(probe_cmd, capture_output=True, text=True, timeout=60, **subprocess_flags())
@@ -684,8 +701,7 @@ def download_video(
ffmpeg_dir = find_ffmpeg_dir() ffmpeg_dir = find_ffmpeg_dir()
if ffmpeg_dir: if ffmpeg_dir:
cmd.extend(["--ffmpeg-location", ffmpeg_dir]) cmd.extend(["--ffmpeg-location", ffmpeg_dir])
if cookies_path and Path(cookies_path).exists(): cmd.extend(_cookie_args(cookies_path))
cmd.extend(["--cookies", cookies_path])
try: try:
with _process_lock: with _process_lock:
+32
View File
@@ -45,6 +45,13 @@ def _bundle_dirs() -> list[Path]:
sub_macos = frameworks / sub / "Contents" / "MacOS" sub_macos = frameworks / sub / "Contents" / "MacOS"
if sub_macos.exists(): if sub_macos.exists():
dirs.append(sub_macos) dirs.append(sub_macos)
else:
# Dev: cerca in <project_root>/bundle_bin/ così `python main.py`
# usa gli stessi binari del bundle (aggiornati via build_*.py)
# invece di Homebrew/PATH che possono essere obsoleti.
project_bundle = Path(__file__).resolve().parent.parent / "bundle_bin"
if project_bundle.exists():
dirs.append(project_bundle)
return dirs return dirs
@@ -145,3 +152,28 @@ def find_ffmpeg() -> Optional[str]:
def find_ffprobe() -> Optional[str]: def find_ffprobe() -> Optional[str]:
"""Ritorna il path completo di ffprobe.""" """Ritorna il path completo di ffprobe."""
return _find_binary("ffprobe") return _find_binary("ffprobe")
def find_fpcalc() -> Optional[str]:
"""Ritorna il path completo di fpcalc (Chromaprint), o None.
Cerca nelle stesse directory di find_ffmpeg / find_ytdlp: bundle
PyInstaller prima, poi percorsi noti (Homebrew su macOS, LOCALAPPDATA
su Windows), infine PATH generico. In dev mode aggiunge anche
`<project_root>/bundle_bin/` così l'app funziona con `python main.py`
dopo aver scaricato fpcalc via build script.
`fpcalc` viene usato dal modulo core.dedup per calcolare fingerprint
audio (identificazione di duplicati).
"""
found = _find_binary("fpcalc")
if found:
return found
# Fallback dev: bundle_bin del progetto
if not _is_frozen():
exe_name = _exe("fpcalc")
project_bundle = Path(__file__).resolve().parent.parent / "bundle_bin"
cand = project_bundle / exe_name
if cand.exists():
return str(cand)
return None
+1 -1
View File
@@ -272,7 +272,7 @@ def update_cover_only(
"--output", str(temp_dir / "cover"), "--output", str(temp_dir / "cover"),
video_url, video_url,
] ]
ffmpeg_dir = _find_ffmpeg_dir() ffmpeg_dir = find_ffmpeg_dir()
if ffmpeg_dir: if ffmpeg_dir:
cmd.extend(["--ffmpeg-location", ffmpeg_dir]) cmd.extend(["--ffmpeg-location", ffmpeg_dir])
if cookies_path and Path(cookies_path).exists(): if cookies_path and Path(cookies_path).exists():
+1
View File
@@ -11,6 +11,7 @@ pythonnet==3.0.5 ; sys_platform == "win32"
clr-loader==0.2.7.post0 ; sys_platform == "win32" clr-loader==0.2.7.post0 ; sys_platform == "win32"
curl_cffi>=0.9.0 curl_cffi>=0.9.0
beautifulsoup4>=4.12.0 beautifulsoup4>=4.12.0
send2trash>=1.8.0
# --- dev only --- # --- dev only ---
pytest>=8.0.0 pytest>=8.0.0
+193
View File
@@ -0,0 +1,193 @@
"""Test per core.dedup — audio fingerprinting via fpcalc + cache SQLite.
Tutti gli unit test usano mock per fpcalc / send2trash: nessuna
integrazione reale, nessun file audio necessario.
"""
from __future__ import annotations
import subprocess
from pathlib import Path
from unittest import mock
import pytest
from core import dedup
# ------------------------------------------------------------------
# Fixture: dedup con cache DB isolato in tmp_path
# ------------------------------------------------------------------
@pytest.fixture
def patched_cache(tmp_path, monkeypatch):
"""Isola la cache SQLite in tmp_path per non toccare il config dir."""
db = tmp_path / "dedup_cache_test.db"
monkeypatch.setattr(dedup, "_cache_db_path", lambda: db)
return db
def _make_fake_audio(tmp_path: Path, name: str, size: int = 1024) -> Path:
"""Crea un file 'audio' fake (byte casuali con estensione .mp3)."""
f = tmp_path / name
f.write_bytes(b"\x00" * size)
return f
# ------------------------------------------------------------------
# scan_folder
# ------------------------------------------------------------------
class TestScanFolder:
def test_no_audio_files(self, tmp_path, patched_cache):
"""Cartella senza audio -> gruppi vuoti, nessuna eccezione."""
# Solo un file .txt (non audio)
(tmp_path / "readme.txt").write_text("hello")
# Anche senza fpcalc disponibile, con 0 audio file ritorna [].
with mock.patch.object(dedup, "find_fpcalc", return_value="/fake/fpcalc"):
groups = dedup.scan_folder(str(tmp_path), recursive=False)
assert groups == []
def test_uses_cache_on_second_scan(self, tmp_path, patched_cache):
"""Prima scan chiama fpcalc; seconda scan riusa la cache."""
_make_fake_audio(tmp_path, "song.mp3", size=2048)
calls: list[str] = []
def fake_compute(fpcalc, path):
calls.append(path)
return {"duration": 180.5, "fingerprint": "FP-A"}
with mock.patch.object(dedup, "find_fpcalc", return_value="/fake/fpcalc"), \
mock.patch.object(dedup, "compute_fingerprint", side_effect=fake_compute), \
mock.patch.object(dedup, "get_bitrate", return_value=320):
# Prima invocazione: fpcalc DEVE essere chiamato
groups1 = dedup.scan_folder(str(tmp_path), recursive=False)
first_calls = len(calls)
# Seconda invocazione (stesso file, stesso mtime/size):
# cache HIT, fpcalc NON viene richiamato
groups2 = dedup.scan_folder(str(tmp_path), recursive=False)
second_calls = len(calls)
assert first_calls == 1, "prima scan deve chiamare fpcalc una volta"
assert second_calls == 1, "seconda scan deve riusare la cache"
# Con un solo file, nessun gruppo di duplicati
assert groups1 == []
assert groups2 == []
def test_groups_duplicates(self, tmp_path, patched_cache):
"""3 file con lo stesso fingerprint -> 1 gruppo di 3, ordinato per bitrate DESC."""
_make_fake_audio(tmp_path, "a.mp3", size=1000)
_make_fake_audio(tmp_path, "b.mp3", size=3000) # size maggiore
_make_fake_audio(tmp_path, "c.mp3", size=2000)
# Tutti stesso fingerprint. Bitrate differente per verificare
# l'ordinamento: b=320 (top), a=192, c=128.
bitrate_by_name = {"a.mp3": 192, "b.mp3": 320, "c.mp3": 128}
with mock.patch.object(dedup, "find_fpcalc", return_value="/fake/fpcalc"), \
mock.patch.object(dedup, "compute_fingerprint",
return_value={"duration": 200, "fingerprint": "SAME-FP"}), \
mock.patch.object(dedup, "get_bitrate",
side_effect=lambda p: bitrate_by_name[Path(p).name]):
groups = dedup.scan_folder(str(tmp_path), recursive=False)
assert len(groups) == 1, "esattamente un gruppo di duplicati"
g = groups[0]
assert len(g) == 3, "tre file nel gruppo"
# Ordine: bitrate DESC -> b (320), a (192), c (128)
assert [Path(e["path"]).name for e in g] == ["b.mp3", "a.mp3", "c.mp3"]
# Ogni entry ha i campi attesi
for e in g:
assert set(e.keys()) >= {"path", "size", "bitrate", "duration", "fingerprint"}
assert e["fingerprint"] == "SAME-FP"
def test_ignores_singletons(self, tmp_path, patched_cache):
"""File con fingerprint unico non compaiono nei gruppi."""
_make_fake_audio(tmp_path, "dup1.mp3")
_make_fake_audio(tmp_path, "dup2.mp3")
_make_fake_audio(tmp_path, "unique.mp3")
fp_by_name = {"dup1.mp3": "FP-X", "dup2.mp3": "FP-X", "unique.mp3": "FP-Y"}
def fake_compute(fpcalc, path):
return {"duration": 100, "fingerprint": fp_by_name[Path(path).name]}
with mock.patch.object(dedup, "find_fpcalc", return_value="/fake/fpcalc"), \
mock.patch.object(dedup, "compute_fingerprint", side_effect=fake_compute), \
mock.patch.object(dedup, "get_bitrate", return_value=256):
groups = dedup.scan_folder(str(tmp_path), recursive=False)
# Solo il gruppo con dup1/dup2
assert len(groups) == 1
names = {Path(e["path"]).name for e in groups[0]}
assert names == {"dup1.mp3", "dup2.mp3"}
# unique.mp3 non appare in nessun gruppo
for g in groups:
for e in g:
assert Path(e["path"]).name != "unique.mp3"
# ------------------------------------------------------------------
# move_to_trash
# ------------------------------------------------------------------
class TestMoveToTrash:
def test_returns_moved_and_failed_summary(self, tmp_path, patched_cache):
"""Verifica che il summary contenga moved/failed correttamente."""
# 2 path OK + 1 path che alza eccezione
ok1 = str(tmp_path / "ok1.mp3")
ok2 = str(tmp_path / "ok2.mp3")
bad = str(tmp_path / "bad.mp3")
def fake_send(p):
if p == bad:
raise OSError("simulated failure")
# ok: no-op
# Il modulo importa send2trash *dentro* la funzione, quindi
# dobbiamo patchare il modulo importato.
with mock.patch("send2trash.send2trash", side_effect=fake_send):
res = dedup.move_to_trash([ok1, bad, ok2])
assert set(res["moved"]) == {ok1, ok2}
assert len(res["failed"]) == 1
assert res["failed"][0]["path"] == bad
assert "simulated failure" in res["failed"][0]["error"]
# ------------------------------------------------------------------
# compute_fingerprint
# ------------------------------------------------------------------
class TestComputeFingerprint:
def test_timeout_returns_error(self):
"""Timeout di fpcalc -> dict con _error, no fingerprint."""
with mock.patch("core.dedup.subprocess.run",
side_effect=subprocess.TimeoutExpired(cmd="fpcalc", timeout=30)):
result = dedup.compute_fingerprint("/fake/fpcalc", "/some/file.mp3")
assert result and "_error" in result and "timeout" in result["_error"]
assert "fingerprint" not in result
def test_success_returns_dict(self):
"""Output JSON valido -> {duration, fingerprint}."""
fake_proc = mock.Mock()
fake_proc.returncode = 0
fake_proc.stdout = '{"duration": 123.4, "fingerprint": "ABCDEF"}'
with mock.patch("core.dedup.subprocess.run", return_value=fake_proc):
result = dedup.compute_fingerprint("/fake/fpcalc", "/some/file.mp3")
assert result == {"duration": 123.4, "fingerprint": "ABCDEF"}
def test_bad_json_returns_error(self):
fake_proc = mock.Mock()
fake_proc.returncode = 0
fake_proc.stdout = "not json at all"
with mock.patch("core.dedup.subprocess.run", return_value=fake_proc):
result = dedup.compute_fingerprint("/fake/fpcalc", "/some/file.mp3")
assert result and "_error" in result and "JSON" in result["_error"]
assert "fingerprint" not in result
def test_missing_fpcalc_returns_error(self):
# Nessuna chiamata subprocess se fpcalc e' vuoto
result = dedup.compute_fingerprint("", "/some/file.mp3")
assert result and "_error" in result and "fpcalc" in result["_error"]
assert "fingerprint" not in result
+144
View File
@@ -1826,3 +1826,147 @@ input[type="number"]::-webkit-inner-spin-button {
word-break: break-all; word-break: break-all;
font-family: "SF Mono", Menlo, Consolas, monospace; font-family: "SF Mono", Menlo, Consolas, monospace;
} }
/* ===============================
DEDUP tab (audio duplicati)
=============================== */
.dedup-summary {
font-size: 14px;
color: var(--text-2);
font-weight: 600;
}
.dedup-group-card {
padding: 14px 16px;
}
.dedup-group-header {
font-size: 13px;
color: var(--text-2);
margin-bottom: 10px;
padding-bottom: 8px;
border-bottom: 1px solid var(--border);
}
.dedup-group-header strong {
color: var(--text-1);
font-size: 14px;
}
.dedup-fp {
font-family: "SF Mono", Menlo, Consolas, monospace;
color: var(--text-3);
font-size: 11px;
}
.dedup-file-list {
display: flex;
flex-direction: column;
gap: 6px;
}
.dedup-file-row {
display: flex;
align-items: center;
gap: 12px;
padding: 8px 10px;
border-radius: 8px;
background: var(--bg-elev);
transition: background 0.15s ease;
}
.dedup-file-row:hover {
background: var(--bg-elev-hover, var(--bg-elev));
}
.dedup-file-check {
flex-shrink: 0;
width: 16px;
height: 16px;
cursor: pointer;
accent-color: var(--red);
}
.dedup-file-check:disabled {
cursor: not-allowed;
opacity: 0.5;
}
.dedup-file-info {
flex: 1;
min-width: 0;
}
.dedup-file-title {
font-size: 13px;
font-weight: 600;
color: var(--text-1);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.dedup-file-meta {
font-size: 11px;
color: var(--text-3);
margin-top: 2px;
font-variant-numeric: tabular-nums;
}
.dedup-file-path {
font-size: 10px;
color: var(--text-3);
margin-top: 2px;
font-family: "SF Mono", Menlo, Consolas, monospace;
word-break: break-all;
opacity: 0.6;
}
.dedup-file-play {
flex-shrink: 0;
display: flex;
align-items: center;
}
/* File "da tenere" (bitrate massimo del gruppo) — verde */
.dedup-keep {
background: var(--green-dim);
border-left: 3px solid var(--green);
}
.dedup-keep .dedup-file-title {
color: var(--green-text);
}
.dedup-keep-badge {
display: inline-block;
padding: 2px 8px;
border-radius: 6px;
background: var(--green);
color: white;
font-size: 10px;
font-weight: 800;
letter-spacing: 0.5px;
margin-right: 6px;
vertical-align: middle;
}
.dedup-footer-bar {
position: sticky;
bottom: 0;
z-index: 5;
display: flex;
justify-content: space-between;
align-items: center;
gap: 12px;
padding: 12px 16px;
margin-top: 12px;
border-radius: var(--r-md);
background: var(--bg-card);
box-shadow: 0 -4px 12px rgba(0, 0, 0, 0.15);
border: 1px solid var(--border);
}
/* Lista risultati Dedup con scroll interno: l'header (hero sticky) e il
footer restano visibili anche se ci sono centinaia di gruppi. */
.dedup-groups-scroll {
max-height: 55vh;
overflow-y: auto;
padding-right: 6px; /* spazio per la scrollbar sul bordo */
}
+98
View File
@@ -49,6 +49,10 @@
<span class="nav-icon">🔄</span> <span class="nav-icon">🔄</span>
<span>Converti</span> <span>Converti</span>
</button> </button>
<button class="nav-item" data-view="dedup" data-feature="audio">
<span class="nav-icon">🗑</span>
<span>Dedup</span>
</button>
<button class="nav-item" data-view="upgrade" data-feature="upgrade"> <button class="nav-item" data-view="upgrade" data-feature="upgrade">
<span class="nav-icon">⚡</span> <span class="nav-icon">⚡</span>
<span>Upgrade</span> <span>Upgrade</span>
@@ -559,6 +563,24 @@
</div> </div>
</div> </div>
<div class="field">
<label class="field-label">Cookies da browser (bypass 403 YouTube)</label>
<div class="row">
<select id="cookiesBrowserSelect" class="input pill" style="max-width:220px;">
<option value="">Nessuno</option>
<option value="chrome">Chrome</option>
<option value="safari">Safari</option>
<option value="firefox">Firefox</option>
<option value="edge">Edge</option>
<option value="brave">Brave</option>
</select>
</div>
<div class="hint">
<span class="hint-ico">ⓘ</span>
<div>Alcuni video YouTube (musica protetta, region-lock, età) danno HTTP 403 senza cookies autenticati. Selezionando un browser, l'app legge i cookies dal tuo profilo — richiede che tu sia loggato su YouTube in quel browser. Se hai anche cookies.txt sopra, ha priorità il file.</div>
</div>
</div>
<div class="field"> <div class="field">
<label class="field-label">Cartella output predefinita</label> <label class="field-label">Cartella output predefinita</label>
<div class="row"> <div class="row">
@@ -1206,6 +1228,82 @@
</div> </div>
</section> </section>
<!-- ===== VIEW: DEDUP (audio duplicati via Chromaprint) ===== -->
<section class="view" id="view-dedup">
<header class="hero hero-red">
<div class="hero-content">
<div class="hero-eyebrow red">DEDUP AUDIO</div>
<h1 class="hero-title">Trova e rimuovi duplicati audio</h1>
<p class="hero-subtitle">Audio fingerprinting via Chromaprint: riconosce brani identici indipendentemente dal bitrate, dal formato o dai tag.</p>
</div>
<div class="hero-deco">🗑</div>
</header>
<h2 class="section-label">Cartella</h2>
<div class="card">
<div class="beatport-header">
<button id="dedup-pick-folder" class="btn btn-primary pill">📁 Scegli cartella</button>
<button id="dedup-start-btn" class="btn btn-primary pill" disabled>
<span class="ico">▶</span> Scansiona
</button>
<button id="dedup-stop-btn" class="btn btn-danger pill" hidden>
<span class="ico">◼</span> Interrompi
</button>
</div>
<div class="beatport-header" style="margin-top:12px;gap:16px;">
<span id="dedup-path-display" class="beatport-output-info" style="margin:0;">Nessuna cartella selezionata</span>
</div>
<div class="beatport-header" style="margin-top:12px;gap:16px;">
<label style="display:flex;gap:6px;align-items:center;color:var(--text-2);font-size:13px;">
<input type="checkbox" id="dedup-recursive" checked />
<span>Ricerca ricorsiva (include sottocartelle)</span>
</label>
</div>
<div class="beatport-header" style="margin-top:12px;gap:16px;">
<label style="display:flex;gap:8px;align-items:center;color:var(--text-2);font-size:13px;">
<span>Metodo:</span>
<select id="dedup-method" class="input" style="padding:4px 8px;font-size:13px;">
<option value="fingerprint">Audio fingerprint (preciso, lento)</option>
<option value="filename">Nome file (veloce, meno preciso)</option>
</select>
</label>
</div>
<div id="dedup-status" class="beatport-status"></div>
</div>
<h2 class="section-label">Progresso</h2>
<div class="card">
<div class="progress">
<div class="progress-bar"><div id="dedupProgressFill" class="progress-fill pink"></div></div>
<div class="progress-meta">
<span id="dedupPercent" class="percent pink">0%</span>
<span id="dedupCounter">In attesa</span>
</div>
</div>
</div>
<h2 class="section-label">Risultati</h2>
<div class="card" id="dedup-summary-card" hidden>
<div class="dedup-summary" id="dedup-summary-line"></div>
</div>
<div id="dedup-groups-wrap" class="dedup-groups-scroll"></div>
<div class="dedup-footer-bar" id="dedup-footer-bar" hidden>
<span class="counter" id="dedup-selected-count">0 file da cancellare</span>
<button id="dedup-select-suggested-btn" class="btn btn-ghost pill">
✓ Seleziona tutti i consigliati
</button>
<button id="dedup-trash-btn" class="btn btn-danger pill" disabled>
<span class="ico">🗑</span> Sposta in cestino
</button>
</div>
<h2 class="section-label">Log</h2>
<div class="card log-card">
<div class="log" id="dedup-log"></div>
</div>
</section>
</main> </main>
</div> </div>
+348
View File
@@ -408,6 +408,7 @@ const logEls = {
youtube: () => $("#youtube-log"), youtube: () => $("#youtube-log"),
convert: () => $("#convert-log"), convert: () => $("#convert-log"),
traxsource: () => $("#traxsource-log"), traxsource: () => $("#traxsource-log"),
dedup: () => $("#dedup-log"),
}; };
function classifyLog(msg) { function classifyLog(msg) {
@@ -484,6 +485,7 @@ async function init() {
$("#bitrateSelect").value = state.config.bitrate || "320K"; $("#bitrateSelect").value = state.config.bitrate || "320K";
$("#hqThresholdInput").value = state.config.hq_threshold || 310; $("#hqThresholdInput").value = state.config.hq_threshold || 310;
$("#cookiesInput").value = state.config.cookies_path || ""; $("#cookiesInput").value = state.config.cookies_path || "";
if ($("#cookiesBrowserSelect")) $("#cookiesBrowserSelect").value = state.config.cookies_browser || "";
$("#outputInput").value = state.config.output_dir || ""; $("#outputInput").value = state.config.output_dir || "";
$("#themeSelect").value = state.config.theme || "dark"; $("#themeSelect").value = state.config.theme || "dark";
applyTheme(state.config.theme); applyTheme(state.config.theme);
@@ -525,6 +527,9 @@ async function init() {
// WAV -> MP3 converter tab // WAV -> MP3 converter tab
await ConvertUI.init(); await ConvertUI.init();
// Dedup tab (audio duplicati via Chromaprint)
await DedupUI.init();
} }
async function refreshRecDevices() { async function refreshRecDevices() {
@@ -1289,6 +1294,7 @@ $("#saveBtn").addEventListener("click", async () => {
bitrate: $("#bitrateSelect").value, bitrate: $("#bitrateSelect").value,
hq_threshold: parseInt($("#hqThresholdInput").value, 10) || 310, hq_threshold: parseInt($("#hqThresholdInput").value, 10) || 310,
cookies_path: $("#cookiesInput").value, cookies_path: $("#cookiesInput").value,
cookies_browser: $("#cookiesBrowserSelect") ? $("#cookiesBrowserSelect").value : "",
output_dir: $("#outputInput").value, output_dir: $("#outputInput").value,
theme: $("#themeSelect").value, theme: $("#themeSelect").value,
}; };
@@ -2857,6 +2863,348 @@ const ConvertUI = (() => {
return { init }; return { init };
})(); })();
// ============================================================
// DedupUI — audio duplicati via Chromaprint fingerprinting
// ============================================================
const DedupUI = (() => {
const mstate = {
folder: "",
recursive: true,
scanning: false,
groups: [],
// Set di path da cancellare (chiave = string path)
toDelete: new Set(),
};
function _fmtDur(sec) {
if (!sec || sec <= 0) return "—";
const m = Math.floor(sec / 60);
const s = Math.floor(sec % 60);
return `${m}:${String(s).padStart(2, "0")}`;
}
function setStatus(text, kind) {
const el = $("#dedup-status");
if (!el) return;
el.textContent = text || "";
el.className = "beatport-status" + (kind ? " " + kind : "");
}
function updateSelectedCount() {
const n = mstate.toDelete.size;
const label = n === 1 ? "1 file da cancellare" : `${n} file da cancellare`;
const cnt = $("#dedup-selected-count");
if (cnt) cnt.textContent = label;
const btn = $("#dedup-trash-btn");
if (btn) btn.disabled = n === 0 || mstate.scanning;
const bar = $("#dedup-footer-bar");
if (bar) bar.hidden = mstate.groups.length === 0;
}
function renderGroups() {
const wrap = $("#dedup-groups-wrap");
if (!wrap) return;
wrap.innerHTML = "";
mstate.toDelete.clear();
if (!mstate.groups.length) {
$("#dedup-summary-card").hidden = true;
$("#dedup-footer-bar").hidden = true;
updateSelectedCount();
return;
}
// Summary line
const nGroups = mstate.groups.length;
let nDupes = 0;
let reclaim = 0;
mstate.groups.forEach((g) => {
nDupes += Math.max(0, g.length - 1);
for (let i = 1; i < g.length; i++) reclaim += (g[i].size || 0);
});
$("#dedup-summary-line").textContent =
`${nGroups} gruppi · ${nDupes} duplicati · spazio recuperabile ~${_fmtBytes(reclaim)}`;
$("#dedup-summary-card").hidden = false;
// Render each group card
mstate.groups.forEach((group, gi) => {
const card = document.createElement("div");
card.className = "card dedup-group-card";
const totalSize = group.reduce((s, e) => s + (e.size || 0), 0);
const fpPreview = (group[0].fingerprint || "").slice(0, 12);
const header = document.createElement("div");
header.className = "dedup-group-header";
header.innerHTML = `
<strong>Gruppo ${gi + 1}</strong>
· <span>${group.length} file</span>
· <span>totale ${_fmtBytes(totalSize)}</span>
· <span class="dedup-fp">fp ${_escapeHtml(fpPreview)}…</span>
`;
card.appendChild(header);
const list = document.createElement("div");
list.className = "dedup-file-list";
group.forEach((entry, idx) => {
const isKeep = idx === 0; // primo = bitrate piu' alto = da tenere
const row = document.createElement("div");
row.className = "dedup-file-row" + (isKeep ? " dedup-keep" : "");
// Pre-selezione: cancella tutti tranne il primo
if (!isKeep) mstate.toDelete.add(entry.path);
const checkbox = document.createElement("input");
checkbox.type = "checkbox";
checkbox.className = "dedup-file-check";
checkbox.checked = !isKeep;
checkbox.disabled = isKeep;
checkbox.title = isKeep
? "File da tenere (bitrate massimo)"
: "Marca per la cancellazione";
checkbox.addEventListener("change", () => {
if (checkbox.checked) mstate.toDelete.add(entry.path);
else mstate.toDelete.delete(entry.path);
updateSelectedCount();
});
row.appendChild(checkbox);
const info = document.createElement("div");
info.className = "dedup-file-info";
const kbps = entry.bitrate ? `${entry.bitrate}k` : "—";
const keepBadge = isKeep ? '<span class="dedup-keep-badge">TIENI</span> ' : "";
info.innerHTML = `
<div class="dedup-file-title">${keepBadge}${_escapeHtml(_basename(entry.path))}</div>
<div class="dedup-file-meta">
${_escapeHtml(kbps)} · ${_escapeHtml(_fmtBytes(entry.size))} · ${_escapeHtml(_fmtDur(entry.duration))}
</div>
<div class="dedup-file-path">${_escapeHtml(entry.path)}</div>
`;
row.appendChild(info);
// Preview audio (riusa _makePreviewBtn della sezione Upgrade)
const playSlot = document.createElement("div");
playSlot.className = "dedup-file-play";
playSlot.appendChild(_makePreviewBtn(entry.path));
row.appendChild(playSlot);
list.appendChild(row);
});
card.appendChild(list);
wrap.appendChild(card);
});
updateSelectedCount();
}
async function pickFolder() {
const p = await window.pywebview.api.dedup_pick_folder();
if (p && typeof p === "string") {
mstate.folder = p;
$("#dedup-path-display").textContent = p;
$("#dedup-start-btn").disabled = false;
setStatus("", "");
}
}
async function startScan() {
if (!mstate.folder) return;
mstate.recursive = $("#dedup-recursive").checked;
mstate.scanning = true;
mstate.groups = [];
if (mstate._groupIndex) mstate._groupIndex.clear();
mstate.toDelete.clear();
renderGroups();
$("#dedup-start-btn").disabled = true;
$("#dedup-pick-folder").disabled = true;
$("#dedup-stop-btn").hidden = false;
$("#dedup-log").innerHTML = "";
$("#dedupProgressFill").style.width = "0%";
$("#dedupPercent").textContent = "0%";
$("#dedupCounter").textContent = "Scansione in corso…";
setStatus("Scansione in corso...", "loading");
const method = $("#dedup-method")?.value || "fingerprint";
let res;
try {
res = await window.pywebview.api.dedup_start_scan({
directory: mstate.folder,
recursive: mstate.recursive,
method,
});
} catch (e) {
setStatus("Errore avvio: " + ((e && e.message) || e), "error");
finishScan();
return;
}
if (!res || !res.ok) {
const msg = (res && res.error) || "Impossibile avviare la scansione";
setStatus(msg, "error");
if (!handleGateBlock(res || {})) toast(msg, "error");
finishScan();
}
}
async function stopScan() {
try { await window.pywebview.api.dedup_stop_scan(); } catch (e) { console.error(e); }
}
function finishScan() {
mstate.scanning = false;
$("#dedup-start-btn").disabled = !mstate.folder;
$("#dedup-pick-folder").disabled = false;
$("#dedup-stop-btn").hidden = true;
updateSelectedCount();
}
async function moveSelectedToTrash() {
const paths = [...mstate.toDelete];
if (!paths.length) return;
const msg = paths.length === 1
? "Spostare 1 file nel cestino?"
: `Spostare ${paths.length} file nel cestino?`;
if (!confirm(msg)) return;
const btn = $("#dedup-trash-btn");
btn.disabled = true;
let res;
try {
res = await window.pywebview.api.dedup_move_to_trash(paths);
} catch (e) {
toast("Errore: " + ((e && e.message) || e), "error");
btn.disabled = false;
return;
}
if (!res || !res.ok) {
toast((res && res.error) || "Errore spostamento", "error");
btn.disabled = false;
return;
}
const moved = new Set(res.moved || []);
// Rimuovi dai gruppi i file spostati; scarta gruppi che scendono sotto i 2
mstate.groups = mstate.groups
.map((g) => g.filter((e) => !moved.has(e.path)))
.filter((g) => g.length >= 2);
renderGroups();
const okN = res.moved_count || 0;
const failN = res.failed_count || 0;
if (failN > 0) {
toast(`${okN} spostati, ${failN} falliti`, "error");
} else {
toast(`${okN} file spostati nel cestino`, "success");
}
}
async function init() {
// Bind eventi
$("#dedup-pick-folder").addEventListener("click", pickFolder);
$("#dedup-start-btn").addEventListener("click", startScan);
$("#dedup-stop-btn").addEventListener("click", stopScan);
$("#dedup-trash-btn").addEventListener("click", moveSelectedToTrash);
$("#dedup-select-suggested-btn").addEventListener("click", () => {
// Reset selezione: per ogni gruppo, marca-per-cancellazione TUTTI
// tranne il primo (il "TIENI" consigliato). Non tocca i primi.
mstate.toDelete.clear();
mstate.groups.forEach((group) => {
for (let i = 1; i < group.length; i++) {
if (group[i] && group[i].path) mstate.toDelete.add(group[i].path);
}
});
// Sync checkbox nel DOM
$$("#dedup-groups-wrap .dedup-file-check").forEach((cb) => {
if (cb.disabled) return; // skip il "TIENI"
cb.checked = true;
});
updateSelectedCount();
});
$("#dedup-recursive").addEventListener("change", (e) => {
mstate.recursive = e.target.checked;
});
// Ripristina ultima cartella + recursive da config
const cfg = state.config || {};
if (cfg.dedup_last_folder) {
mstate.folder = cfg.dedup_last_folder;
$("#dedup-path-display").textContent = cfg.dedup_last_folder;
$("#dedup-start-btn").disabled = false;
}
if (typeof cfg.dedup_recursive === "boolean") {
$("#dedup-recursive").checked = cfg.dedup_recursive;
mstate.recursive = cfg.dedup_recursive;
}
if (cfg.dedup_method && $("#dedup-method")) {
$("#dedup-method").value = cfg.dedup_method;
}
// Bridge handlers
bridgeHandlers["dedup:progress"] = (p) => {
if (!p) return;
if (typeof p.overall === "number") {
const pct = Math.round(p.overall * 100);
$("#dedupProgressFill").style.width = pct + "%";
$("#dedupPercent").textContent = pct + "%";
}
if (typeof p.idx === "number" && typeof p.total === "number" && p.total > 0) {
const status = p.status || "";
let label;
if (status === "cached") label = `File ${p.idx}/${p.total} · cache`;
else if (status === "computing") label = `File ${p.idx}/${p.total} · calcolo`;
else if (status === "completed") label = `Completato: ${p.total} file`;
else if (status === "stopped") label = "Interrotta";
else if (status === "error") label = `File ${p.idx}/${p.total} · errore`;
else label = `File ${p.idx}/${p.total}`;
$("#dedupCounter").textContent = label;
}
};
// Streaming: gruppi arrivano uno alla volta man mano che si formano.
// mstate.groups è una lista di ARRAY di entries (shape richiesta da
// renderGroups). Manteniamo un indice group_id → arrayRef per poter
// aggiornare in-place quando il gruppo cresce.
if (!mstate._groupIndex) mstate._groupIndex = new Map();
bridgeHandlers["dedup:group"] = (p) => {
if (!p || !p.id || !Array.isArray(p.entries)) return;
const idx = mstate._groupIndex;
const existing = idx.get(p.id);
if (existing) {
// Aggiorna elementi in-place (stesso riferimento array in mstate.groups)
existing.length = 0;
for (const e of p.entries) existing.push(e);
} else {
// Nuovo gruppo: array copia + tag `_id` non-enumerabile per il lookup
const arr = p.entries.slice();
Object.defineProperty(arr, "_id", { value: p.id, enumerable: false });
idx.set(p.id, arr);
mstate.groups.push(arr);
}
renderGroups();
setStatus(`Trovati ${mstate.groups.length} gruppi finora…`, "loading");
};
bridgeHandlers["dedup:done"] = (p) => {
// Se sono arrivati group via streaming, ignoriamo p.groups (già in mstate).
// Fallback: se lo streaming non ha popolato nulla, usiamo p.groups.
if (mstate.groups.length === 0 && p && Array.isArray(p.groups)) {
mstate.groups = p.groups;
}
renderGroups();
if (p && p.ok === false) {
setStatus(p.error || "Errore scansione", "error");
} else if (!mstate.groups.length) {
setStatus("Nessun duplicato trovato", "ok");
} else {
setStatus(`Trovati ${mstate.groups.length} gruppi di duplicati`, "ok");
}
finishScan();
};
}
return { init };
})();
// ============================================================ // ============================================================
// Boot // Boot
// ============================================================ // ============================================================