Initial commit

This commit is contained in:
jkortis committed 2026-09-30 23:06:14 -04:00
commit 1d235d30e7
58 files changed
+19693

No files matched your search

+31
View File
@@ -0,0 +1,31 @@
# =====================================================================
# Media Sorter - Configuration Template (.env.example)
# =====================================================================
# 1. Source directory: folder where new downloads arrive
DOWNLOADS_DIR=./downloads
# 2. Destination directories: where media files are sorted to
MOVIES_DIR=./movies
SHOWS_DIR=./shows
# 3. Execution mode:
# Set to 'false' to immediately move files into movies/shows.
# Set to 'true' to safely preview operations without touching files.
DRY_RUN=false
# Action type: move | copy | link | hardlink
ACTION=move
# Confidence score required (0.0 to 1.0). Uncertain files go to Quarantine review
CONFIDENCE_THRESHOLD=0.75
# Minimum file age in seconds before processing (0 = immediate sorting)
MIN_FILE_AGE_SECONDS=0
# 4. Web Dashboard Server
SERVER_HOST=0.0.0.0
SERVER_PORT=8085
# 5. Database location
DATABASE_PATH=media_sorter.db
+54
View File
@@ -0,0 +1,54 @@
# Byte-compiled / optimized / DLL files
__pycache__/
*.py[cod]
*$py.class
# Caches and test artifacts
.pytest_cache/
.hypothesis/
.coverage
htmlcov/
*.log
# Environments
.venv/
env/
venv/
ENV/
env.bak/
venv.bak/
# Local database & runtime state
*.db
*.db-shm
*.db-wal
*.lock
media_sorter.db
media_sorter.lock
media_sorter_posters.json
# Environment variables (contains private user paths)
.env
# Runtime sample reports & local media directories
sample_dry_run_report.json
downloads/
movies/
shows/
incoming/
organized/
# OS generated files
.DS_Store
.DS_Store?
._*
.Spotlight-V100
.Trashes
ehthumbs.db
Thumbs.db
# User review images
prompt_images/
# Agent working directories
.agents/
+240
View File
@@ -0,0 +1,240 @@
<div align="center">
# 🎬 Media Sorter
**High-performance, automated media classification and organization engine with defensive data safety, atomic operations, transactional rollback, and modern web UI.**
[![Release](https://img.shields.io/badge/release-v1.0.0-blue.svg)](https://github.com/)
[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)
[![FastAPI](https://img.shields.io/badge/FastAPI-0.100+-009688.svg)](https://fastapi.tiangolo.com)
[![SQLite WAL](https://img.shields.io/badge/SQLite-WAL%20Journaling-003B57.svg)](https://www.sqlite.org/wal.html)
[![Tests Passing](https://img.shields.io/badge/tests-51%20passing-brightgreen.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
</div>
---
## 🌟 Highlights
- 🛡️ **Zero Silent Data Loss Guarantee**: Files are never deleted or silently overwritten. Unsure files are safely quarantined.
- 🔍 **Multi-Signal Classification Engine**: Classifies media using filename patterns, tokens, codecs, container streams, ID3v2/Vorbis tags, and episode hierarchies.
- 🎨 **Modern Interactive Web UI**: Fully responsive dashboard with customizable themes, Folder Explorer with show detection & live TV show artwork, Quarantine Queue, and Settings.
- ⚡ **Non-Shifting Scrollable Viewport**: Smoothly scroll through 500+ file libraries with sticky table headers without shifting the top navigation bar.
- ↩️ **Atomic Operations & 1-Click Rollback**: Every batch is journaled in SQLite WAL mode. Easily invert any batch with a single click or CLI command.
- 🚀 **Production Ready**: Native PM2 integration, Docker & Docker Compose setup, and systemd service templates included.
---
## 📑 Table of Contents
- [Supported Media Types](#-supported-media-types)
- [Safety & Reliability Principles](#-safety--reliability-principles)
- [Web Dashboard Features](#-web-dashboard-features)
- [Quickstart & Installation](#-quickstart--installation)
- [Option A: Python Virtualenv](#option-a-python-virtualenv)
- [Option B: PM2 Process Manager](#option-b-pm2-process-manager-recommended-for-servers)
- [Option C: Docker & Docker Compose](#option-c-docker--docker-compose)
- [Configuration (.env)](#-configuration-env)
- [CLI Reference](#-cli-reference)
- [REST API Reference](#-rest-api-reference)
- [Testing & Quality Assurance](#-testing--quality-assurance)
- [License](#-license)
---
## 🎯 Supported Media Types
| Category | Typical Formats | Detection Signals |
| :--- | :--- | :--- |
| **Movies** | `.mkv`, `.mp4`, `.avi`, `.m4v` | Year tags, edition flags (Extended/Director's Cut), resolution tokens, single file structures. |
| **TV Shows** | `.mkv`, `.mp4`, `.ts` | `S01E05`, `1x05`, season pack structures, multi-episode tokens, episode titles. |
| **Anime** | `.mkv`, `.mp4` | Fansub release brackets `[SubsPlease]`, absolute numbering (`Episode 500`), CRC32 hashes. |
| **Music** | `.flac`, `.mp3`, `.m4a`, `.opus` | ID3v2/Vorbis metadata, disc/track tags, multi-disc hierarchies, artist tokens. |
| **Audiobooks** | `.m4b`, `.mp3` | Chapter tags, narrator metadata, audiobook series tags. |
| **Podcasts** | `.mp3`, `.m4a` | Release dates (`YYYY-MM-DD`), episode numbers, show titles. |
| **Documentaries**| `.mkv`, `.mp4` | Miniseries tags, broadcast metadata, documentary tokens. |
| **Sidecars** | `.srt`, `.ass`, `.nfo`, images | Automatically mapped and moved alongside parent media files. |
| **Quarantine** | Any unclassified / low confidence | Isolated safely in the quarantine queue for user review. |
---
## 🛡️ Safety & Reliability Principles
1. **Dry-Run by Default**: All actions run in simulation mode unless explicitly triggered as Live.
2. **Confidence Thresholding**: Items with classification score `< 0.75` (configurable) are diverted to Quarantine rather than misplaced.
3. **Collision Avoidance**: If a target file already exists, Media Sorter executes your collision policy (`rename_unique`, `replace_if_higher_quality`, `quarantine`, `skip`, or `error`).
4. **Atomic Two-Stage Moves**: Same-filesystem operations use atomic `os.replace`. Cross-filesystem operations write to hidden temporary files, verify size and integrity, atomically link, and only then unlink the source.
5. **Active File Protection**: Files actively being written by BitTorrent or downloading clients (or modified within `min_file_age_seconds`) are skipped until finished.
6. **Transactional Journaling**: Every operation logs old path, new path, file size, hash, and status into a SQLite database with WAL journaling.
---
## 🖥️ Web Dashboard Features
### 1. Folder Explorer with Live Show Artwork & Accordions
- **Intelligent Grouping**: Automatically identifies episodes belonging to the same series in your downloads folder.
- **Show Artwork**: Displays official high-resolution posters from TVmaze and local directories next to the believed show name.
- **Collapsible Cards**: Expand and collapse individual shows or use "Expand All" / "Collapse All" controls.
- **Single Files Table**: Non-episodic files and movies are cleanly separated with detected metadata.
### 2. Dedicated Settings Modal & Process Control
- Access settings via the **⚙️ Settings** button in the header.
- Switch visual themes with live color cards.
- **Restart Server Process**: Restarts the PM2 process with an automatic reconnection overlay.
- **Clear Activity History**: Wipes historical batch records with a single click.
### 3. Quarantine Review & Resolution
- Review items with lower confidence.
- Sort them into **Movies** or **Shows** with one click.
---
## 🚀 Quickstart & Installation
### Option A: Python Virtualenv
```bash
# 1. Clone repository
git clone https://github.com/yourusername/media-sorter.git
cd media-sorter
# 2. Create and activate virtual environment
python3 -m venv .venv
source .venv/bin/activate
# 3. Install package and dependencies
pip install -e .
# 4. Copy and configure .env
cp .env.example .env
nano .env
# 5. Start the web dashboard
media-sorter web --host 0.0.0.0 --port 8085
```
### Option B: PM2 Process Manager (Recommended for Servers)
```bash
# 1. Install PM2 globally (if not already installed)
npm install -g pm2
# 2. Start Media Sorter using ecosystem.config.js
pm2 start ecosystem.config.js
# 3. Save PM2 startup list
pm2 save
pm2 startup
```
To view logs or restart:
```bash
pm2 logs media-sorter
pm2 restart media-sorter --update-env
```
### Option C: Docker & Docker Compose
```bash
# Configure paths in docker-compose.yml or .env, then launch:
docker compose up -d
```
---
## ⚙️ Configuration (.env)
Create a `.env` file in the project root:
```ini
# Directories (Absolute or Relative Paths)
DOWNLOADS_DIR=/path/to/downloads # Incoming media source
MOVIES_DIR=/path/to/movies # Destination for movies
SHOWS_DIR=/path/to/tv # Destination for TV shows and anime
# Operational Mode
DRY_RUN=false # true = preview only, false = perform actual moves
CONFIDENCE_THRESHOLD=0.75 # Minimum confidence to automatically sort (0.0 - 1.0)
ACTION=move # 'move', 'copy', or 'hardlink'
SCAN_INTERVAL=0 # Background daemon scan interval in seconds (0 = disabled)
MIN_FILE_AGE=0 # Minimum file age in seconds before processing
# Web Server
PORT=8085
HOST=0.0.0.0
```
---
## 💻 CLI Reference
Media Sorter includes a full CLI for scripting and headless servers:
```bash
# Run a dry-run preview on downloads
media-sorter run --dry-run
# Run live sorting
media-sorter run --live
# Rollback a specific batch
media-sorter rollback --batch-id <BATCH_UUID>
# Rollback the most recent batch
media-sorter rollback --latest
# Review quarantined files
media-sorter quarantine list
# Launch web dashboard
media-sorter web --port 8085
```
---
## 🌐 REST API Reference
The built-in FastAPI server provides endpoints for dashboard integrations:
| Method | Endpoint | Description |
| :--- | :--- | :--- |
| `GET` | `/` | Serves the interactive web interface. |
| `GET` | `/api/status` | Current server configuration, paths, and stats. |
| `GET` | `/api/files` | Discovered files, show groupings, and poster URLs. |
| `POST` | `/api/run` | Execute sort run (`{"dry_run": true/false}`). |
| `POST` | `/api/rollback` | Rollback a previous batch (`{"batch_id": "..."}`). |
| `GET` | `/api/batches` | List batch history and statuses. |
| `POST` | `/api/batches/clear` | Clear batch and activity history. |
| `GET` | `/api/quarantine` | List files currently held in quarantine. |
| `POST` | `/api/quarantine/resolve`| Manually resolve a quarantined item (`movie` or `tv`). |
| `GET` | `/api/poster` | Query or fetch show poster artwork URL. |
| `POST` | `/api/settings` | Save updated `.env` configuration. |
| `POST` | `/api/restart` | Gracefully restart the server process. |
---
## 🧪 Testing & Quality Assurance
Media Sorter includes a comprehensive suite of 51 unit, integration, and fuzz tests:
```bash
# Run tests
pytest
# Run tests with verbose output
pytest -v
```
Test coverage includes:
- Multi-token classification and fuzzy filename parsing.
- High-concurrency database journaling in SQLite WAL mode.
- Cross-filesystem atomic transfer simulation.
- 100% rollback fidelity across complex batches.
- Web API endpoints and environment reconfiguration.
---
## 📄 License
This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
+181
View File
@@ -0,0 +1,181 @@
# Changelog
All notable changes to the Media Sorter project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [1.1.0] - 2026-09-07
### Added
- **Robust Long-Running Anime Tokenization & Season/Episode Extraction**:
- Full support for 4-digit episode numbers (`\d{1,4}`) across standard TV tags (`S01E1171`) and standalone episode prefixes (`EP1171`, `E233`).
- Added support for anime releases with episode titles (`[Anime Time] One Piece - 0233 - Episode Title [Tags]`), space-separated formats (`[HatSubs] One Piece 1073...`), and underscore formats (`[Kaerizaki-Fansub]_One_Piece_1076...`).
- Parent directory season extraction (`Season 08 - Water Seven`) scanning ancestor paths from closest to farthest while ignoring system root folders.
- **Natural Ordering & Web UI Auto-Detection**:
- Implemented `natural_sort_key` in Python and `naturalCompare` in JavaScript so episodes never get misaligned or sequentially numbered out of order.
- Added "Auto-Detect from Filenames" button in the manual batch modal to immediately pull true season and episode numbers directly from file names.
- Increased episode numbering limit to 9999 to support long-running anime like *One Piece*.
- Fixed `clean_detected_show_name` to preserve year-like numbers (1900-2099) in movie and series titles (e.g. *Blade Runner 2049*).
- **Manual Batch Sorting & Library Management**:
- Added `POST /api/files/manual-sort-batch` and `POST /api/library/clear` endpoints.
- Added floating multi-file selection bar and 2-step manual batch sorting overlay modal.
## [1.0.7] - 2026-09-05
### Added
- **Library Catalog & Show/Movie Management (`📚 Library`)**:
- Added a dedicated top-level **📚 Library** view with full catalog tracking of all indexed shows and movies across storage drives.
- Interactive selector buttons to switch between **📺 Shows**, **🎬 Movies**, and **📂 All**, with live count badges.
- Real-time instant search input to filter titles as you type.
- Disk scanning synchronization (`POST /api/library/rescan`) that indexes directory trees, episode counts, seasons, and technical metadata.
- Persistent database tracking via `LibraryItem` table model with unique constraint protection and last-updated tracking.
- **Automatic Show Memory Routing**:
- Whenever a new show is detected or sorted, it is automatically cataloged in the Library.
- When inspecting downloads or analyzing files, media sorter queries known shows from the library and automatically routes incoming media (such as extras, behind-the-scenes, specials, and unbracketed files) into that show's folder (`SHOWS_DIR / <Show Name>`).
- **Unsure File Grouping Dropdowns**:
- Automatically clusters non-show files into dedicated collapsible dropdown cards instead of dumping them into a flat singles list:
- **Shared Subfolder Groups**: Files sharing a common subfolder in downloads form their own group card (e.g. `📁 Shared Subfolder: Dexter Extras (5 files)`).
- **Matching Name Prefix Groups**: Files sharing a common clean title stem or prefix form their own group card (e.g. `🏷️ Matching Name Prefix: Blood, Guts and Body Parts (2 files)`).
- Only truly lone, isolated files remain in the singles media list.
- Group action buttons:
- `⚡ Sort Group`: Sorts all files in the group directly via `POST /api/files/sort-group`.
- `✏️ Set Show/Movie for Group`: Opens the manual assignment modal for the entire group, allowing one-click bulk categorization to Movies or TV Shows.
- **Group Sorting Endpoint (`POST /api/files/sort-group`)**:
- Server endpoint that batch-processes unsure groups of files with show name override or movie destination formatting, atomic movement, and automatic library catalog registration.
### Fixed
- **Anime Fansub Adjacent Brackets Regex**:
- Fixed `RE_ANIME_RELEASE` regex to support adjacent bracket tags without spaces (e.g. `[Erai-raws] Bleach - 001 [1080p][MultiSub][EF0AF7BA].mkv`), ensuring anime episode fan releases are accurately classified.
- **Prevent Anime Episode Matching into SxxExx**:
- Added negative lookahead `(?![xX\w])` to prevent episodic titles like `Game of Thrones - 1x09 - Baelor` from prematurely matching `1` as an anime episode instead of `Season 1 Episode 9`.
- **Fast Local Artwork & Responsive Downloads Inspection**:
- Updated `fetch_show_poster` with `allow_network=False` during folder inspection so downloads scanning completes in under 2 seconds rather than stalling on synchronous external network queries.
## [1.0.6] - 2026-09-05
### Added
- **Targeted Show Sorting (`POST /api/files/sort-show`)**:
- The "⚡ Sort Now" button inside any show's dropdown card in the Downloads Folder Explorer now targets that specific show directly.
- Sorts all episodes of that show into `SHOWS_DIR / <Show Name> / Season XX / ...` with full transaction tracking, database history, and instant rollback support.
- Interactive UI with button loading spinner (`<span class="loading-spinner"></span> Sorting...`), disabled state during execution, success toast with count of sorted episodes, and automatic UI refresh.
- **Standalone Episode Pattern Recognition**:
- Added support for standalone `Episode \d{1,4}` and `Ep \d{1,4}` patterns (e.g., `Naruto Episode 207 The Supposed Sealed Ability.mkv`, `Bleach Episode 05.mkv`).
- Added support for anime/fansub series formatting without release group prefixes (e.g., `BLEACH꞉ Sennen Kessen-hen - 27 E89717B7].mkv`).
- Added support for anime opening/ending patterns (`S03ED01`, `S03OP01`).
- **Interactive Button Loading States**:
- Added loading spinner and disabled state to `⚡ Run Sort Now (Live)`, `⚡ Sort All Files`, and show-specific `⚡ Sort Now` buttons to prevent duplicate runs and give immediate visual feedback.
### Fixed
- **Show "Sort Now" Button Doing Nothing**:
- Fixed issue where clicking "Sort Now" inside a show dropdown appeared to do nothing because episodic shows like Naruto were failing tokenizer matching, getting classified as low-confidence movies, and being quarantined (which were flagged but not moved).
- Fixed button calling global sorting instead of targeted show sorting.
- Fixed run results panel being hidden on Dashboard tab when viewing Folder Explorer tab.
- **MKV Video Containers Misclassified as Audio/Music**:
- Fixed bug where video containers (`.mkv`, `.mp4`) containing audio streams were routed to audio-only classification if lightweight EBML header parsing didn't match a hardcoded video codec string.
- Reordered classifier pipeline so video containers and extensions are always checked as videos first before audio-only handling.
- Expanded EBML video track detection in `_parse_ebml` to recognize generic `V_` track headers (including VC-1, DivX, XviD, and MPEG2).
## [1.0.5] - 2026-09-05
### Added
- **Quarantine Top Action Buttons & Bulk Controls**:
- Action buttons at the top of the Review & Quarantine header for immediate manual actions: `🎬 Move to Movies`, `📺 Move to Shows`, `✏️ Set Show/Movie`, `↩️ Unflag`, and `↩️ Undo All Resolved`.
- Multi-item selection checkboxes with header `Select All` / `Deselect All` toggle and live selection count tracking.
- Selection toolbar that adapts dynamically based on checked items for bulk execution.
- Backend bulk endpoints:
- `POST /api/quarantine/bulk-resolve`: Bulk moves selected or all pending quarantine items to Movies or TV Shows inside a fast transaction.
- `POST /api/quarantine/bulk-undo`: Bulk unflags pending items or rollbacks all resolved quarantine files back to original source directories.
- **Manual Modal File Switcher**:
- Added file picker dropdown inside `modal-manual-sort` when multiple quarantine items are available, enabling switching between items without reopening the modal.
- **Automatic Folder Explorer & Queue Refresh**:
- Automatically re-scans and refreshes the Downloads Folder Explorer (`loadFiles`) and Quarantine queue (`loadQuarantine`) immediately whenever sorting runs (live or preview) or batch rollbacks complete.
- **Granular & Global Undo Actions**:
- Added `↩️ Undo All Resolved` button to easily restore all resolved quarantine files to their source folders.
- Per-item unflagging for pending items and per-item rollback for resolved items.
### Fixed
- **False-Positive TV Show Classifications for Movies**:
- **Resolution Dimensions**: Filenames containing dimensions like `1920x1080` or `1280x720` no longer falsely match episodic TV season/episode regex (`20x108`, `80x720`). Added negative lookbehind/lookahead and resolution fallback parser `RE_DIMENSIONS`.
- **Release Years in Torrent Tags**: Anime release regex pattern no longer treats 4-digit release years (1900–2099) in bracketed torrent tags (e.g., `[YTS.MX] Movie Title - 2024 [1080p].mkv`) as episode numbers.
- **Broad Substring Match**: Replaced broad `"tv" in path_str` check with directory boundary regex `(?i)[/\\](?:tv[/\\]|tv[-_\s]shows?|tv[-_\s]series|season[-_\s]*\d+)`, preventing movies with `HDTV`, `[rartv]`, or `Apple.TV` from routing to TV show logic.
- **Folder Explorer Subdirectories**: Fixed downloads inspector so scene/torrent movie subfolders (e.g. `The.Dark.Knight.2008.1080p.BluRay/`) are not misidentified as TV show names; movies now properly appear under `singles` targeting `MOVIES_DIR`.
- **Release Group Cleaning**: Updated `_clean_title` to cleanly strip leading bracketed group tags (e.g., `[YTS.MX]`, `[rarbg]`, `[TGx]`).
- **`.txt` Companion Cleaning & Explorer Filter**:
- Excluded `.txt` companions (e.g., torrent notes, RARBG.txt) from the folder explorer.
- Automatically cleans up companion `.txt` files and empty parent release folders when media files are deleted or organized.
- **Safe Quarantine Mode**:
- Flagged files with confidence below threshold remain in place safely rather than being automatically moved.
- **Missing Module Import**:
- Fixed missing `import re` in `src/media_sorter/classifier.py`.
### Changed
- Unified Settings and `.env` configuration into a single coherent dashboard tab.
- Expanded automated test suite to 65 tests covering movie dimension tokenization, bracketed release years, HDTV classification, folder explorer movie detection, and bulk quarantine APIs.
---
## [1.0.4] - 2026-09-05
### Added
- File renaming toggle (`RENAME_FILES`) and customizable naming templates for movies and TV shows (`MOVIE_TEMPLATE`, `TV_TEMPLATE`).
- Automatic cleanup of empty parent directories after files are moved (`CLEANUP_EMPTY_DIRS`).
---
## [1.0.3] - 2026-09-05
### Added
- TV show dropdown aggregation in downloads folder explorer with collapsible season groups.
- Local artwork and TVmaze poster lookup for series.
---
## [1.0.2] - 2026-09-05
### Added
- Instant batch rollbacks and transactional operation journaling.
- Manual file categorization modal and REST API endpoints.
---
## [1.0.1] - 2026-09-05
### Added
- Multi-signal classifier supporting audiobooks, podcasts, home videos, and archives.
- SQLite WAL operation journaling.
---
## [1.0.0] - 2026-09-05
### Added
- Initial release with atomic file sorting, dry-run simulation, and web dashboard.
## [1.0.8] - 2026-09-06
### Added
- Added missing custom themes: neon-forest, retro-retro, golden-sand, deep-space, candy-cotton.
- Implemented RGB Chroma mode UI with toggle and speed slider, keyframe animations, and persistence via localStorage.
- Added unit tests for theme availability and RGB feature.
## [1.0.9] - 2026-09-06
### Added
- Implemented `.rgb-mode` CSS class with keyframe animation for chroma glow.
- Updated project version to 1.0.9 across metadata.
### Fixed
- RGB visual effect was missing, now correctly applied.
- Version numbers were out‑of‑sync between `pyproject.toml` and package `__init__`.
## [1.0.10] - 2026-09-06
### Added
- Updated README with section on using the Antigravity bot via Direct Messages.
- Added no‑cache HTTP headers to ensure version badge updates reliably.
### Changed
- Bumped project version to 1.0.10 in `pyproject.toml` and package `__init__`.
### Fixed
- Version numbers now stay in sync across metadata files.
+67
View File
@@ -0,0 +1,67 @@
# =====================================================================
# Media Sorter - Production Multi-Stage Dockerfile
# =====================================================================
FROM python:3.12-slim AS builder
WORKDIR /build
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc \
build-essential \
&& rm -rf /var/lib/apt/lists/*
COPY pyproject.toml README.md ./
COPY src/ ./src/
RUN pip install --no-cache-dir --upgrade pip && \
pip install --no-cache-dir build && \
python -m build --wheel
# ---------------------------------------------------------------------
# Final Production Runtime Image
# ---------------------------------------------------------------------
FROM python:3.12-slim AS runner
LABEL org.opencontainers.image.title="Media Sorter" \
org.opencontainers.image.description="Reliable, high-performance media classification and organization system" \
org.opencontainers.image.licenses="MIT"
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
curl \
mediainfo \
ffmpeg \
gosu \
&& rm -rf /var/lib/apt/lists/*
# Create non-privileged user and group
RUN groupadd -g 1000 mediasorter && \
useradd -u 1000 -g mediasorter -m -s /bin/bash mediasorter && \
mkdir -p /media/incoming /media/organized /config /data && \
chown -R mediasorter:mediasorter /media /config /data /app
# Copy wheel from builder and install
COPY --from=builder /build/dist/*.whl /tmp/
RUN pip install --no-cache-dir /tmp/*.whl && rm -f /tmp/*.whl
# Default environment variables
ENV PYTHONUNBUFFERED=1 \
MEDIA_SORTER_DATABASE__PATH=/data/media_sorter.db \
MEDIA_SORTER_STORAGE__SOURCE_DIRS='["/media/incoming"]' \
MEDIA_SORTER_STORAGE__DESTINATION_BASE=/media/organized \
MEDIA_SORTER_SERVER__HOST=0.0.0.0 \
MEDIA_SORTER_SERVER__PORT=8080
EXPOSE 8080
# Healthcheck testing the REST status endpoint
HEALTHCHECK --interval=30s --timeout=5s --start-period=5s --retries=3 \
CMD curl -f http://localhost:8080/api/status || exit 1
# Run as non-root user
USER mediasorter
ENTRYPOINT ["media-sorter"]
CMD ["server", "--host", "0.0.0.0", "--port", "8080", "--config", "/config/config.yaml"]
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Media Sorter Contributors
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+162
View File
@@ -0,0 +1,162 @@
# Original User Request
## Initial Request — 2026-09-05T18:48:23Z
Conduct a comprehensive codebase audit, architectural decoupling, database transaction hardening, and test suite expansion for the Media Sorter application to ensure production-grade reliability, modularity, and maintainability.
Working directory: /md0/media-sorter
Integrity mode: development
## Requirements
### R1. Architectural Decoupling and Modularity
Audit the codebase and decompose monolithic components (notably server orchestration, UI template rendering, and endpoint handling) into clean, decoupled modules with clear separation of concerns, while maintaining strict backward compatibility with the existing single-page web dashboard and REST API contract.
### R2. Database Reliability and Transactional Safety
Harden database session management, connection pooling, and error handling across concurrent requests, ensuring clean lifecycle scoping (preventing generator/context manager misuse, uncommitted mutations, or leaked connections) and robust transaction rollback resilience.
### R3. Error Handling and Input Validation
Harden all API endpoints and sorting pipelines with rigorous data validation, clear structured error reporting, and defensive fallbacks to handle malformed inputs, unreachable paths, and unexpected filesystem states gracefully without crashing the service.
### R4. Test Suite Expansion and Edge-Case Coverage
Expand the automated test suite with extensive unit, property-based, and edge-case integration tests covering complex media filenames (anime release tags, roman numerals, multi-part episodes, ambiguous years), concurrent sort operations, and mock filesystem isolation.
### R5. Controlled Infrastructure and Filesystem Safety Guardrails
All file manipulation and verification tests must execute within isolated temporary test fixtures. The system must never mutate or delete real user media files in production directories (`/md0/jdownloads`, `/md0/movies1`, `/md0/tv1`) during testing or audit execution.
## Acceptance Criteria
### Test Verification & Quality
- [ ] `pytest tests/` runs with a 100% pass rate across all existing 75 unit/integration tests and any newly added tests with zero regressions.
- [ ] Core sorting, tokenization, database, and API routing modules achieve or exceed 90% test coverage.
- [ ] All new and modified tests use isolated mock directories or temporary fixtures without touching real user media paths.
- [ ] Static syntax, import integrity, and code quality checks pass cleanly.
### Functional & Contract Compatibility
- [ ] All existing REST API endpoints (`/api/status`, `/api/files`, `/api/run`, `/api/rollback`, `/api/library`, `/api/quarantine`, etc.) remain 100% backward-compatible in request signatures and response schemas.
- [ ] Single-page web dashboard remains fully functional, retaining Folder Explorer (with `.txt`/`.srt` exclusions), Library catalog views, quarantine queue, theme dropdown & swatches, and manual classification modals.
- [ ] Existing companion file behaviors (subtitle pairing and cleanup routines) function without regressions.
## Follow-up — 2026-09-05T23:48:06Z
Expand Media Sorter's automated media sorting rules, classification edge-case handling (TV series, Anime, Movies, multi-part episodes, specials, year tags), and build a comprehensive automated test suite and regression benchmark.
Working directory: /md0/media-sorter
Integrity mode: development
## Requirements
### R1. Robust Pattern Matching & Categorization
Expand media classification and destination routing to reliably distinguish between:
- Movies (including release years, editions, multi-part movies)
- TV Shows (standard SxxExx, season packs, specials, daily/dated shows)
- Anime formats (absolute episode numbers, batch tags, resolution tokens)
Ensure filenames with resolution tags (e.g. 1080p, 4K), audio codecs, release groups, and complex punctuation are accurately normalized without losing episode or title details.
### R2. Isolated Automated Verification Suite
Develop a repeatable test suite testing classification, parsing, and destination resolution against a diverse benchmark dataset of real-world media filename patterns.
- Tests must run offline without requiring external network connectivity or mutations to real media storage folders.
- Provide clear reporting on passed, failed, and edge-case classifications.
### R3. Safe Backward Compatibility
Ensure existing UI operations and server API endpoints (`/api/files/scan`, `/api/files/sort-show`, `/api/library`) continue functioning seamlessly with any pattern parser enhancements.
## Acceptance Criteria
### Test Execution
- [ ] An automated test suite command (e.g., `pytest` or a test runner script) executes cleanly within the workspace.
- [ ] 100% of benchmark test cases pass with zero unhandled exceptions.
- [ ] No regression on existing sorting endpoints or file categorization routes.
## Follow-up — 2026-09-06T02:34:04Z
Conduct a comprehensive feature enhancement, architectural decoupling, and thorough automated bug testing suite execution for the Media Sorter application.
Working directory: /md0/media-sorter
Integrity mode: development
## Requirements
### R1. Architectural Decoupling & Server Modularity
Decompose monolithic components (server routing, API endpoints, file operations) into decoupled, maintainable modules while preserving complete backward compatibility with the existing single-page web dashboard and REST API contract.
### R2. Robust Media Classification & Edge-Case Sorting
Enhance pattern matching and metadata extraction to accurately categorize complex real-world media:
- Standard Movies and TV series (including season/episode tags, resolutions, and audio codecs)
- Anime releases (absolute episode numbering, release group tags)
- Multi-part episodes, specials, and dated daily shows
Ensure companion files (subtitles .srt, .sub) and folder cleanups function reliably.
### R3. Resilient Database & Transaction Scoping
Harden SQLite WAL database operations, batch history tracking, and rollback mechanics to guarantee transactional integrity and prevent data corruption during concurrent operations.
### R4. Comprehensive Bug Testing & Verification Suite
Audit and test the entire application pipeline against edge-case filenames, malformed inputs, unreachable paths, and concurrent requests with isolated mock fixtures.
## Acceptance Criteria
### Verification & Bug Testing
- [ ] Automated test suite (`pytest tests/`) passes 100% across all unit, integration, and safety tests with zero unhandled exceptions.
- [ ] Media classification and tokenizer handle standard, anime, movie, special, and messy filename patterns accurately.
- [ ] Rollback mechanics correctly restore moved files across individual batches and full batch histories.
- [ ] Filesystem safety guardrails ensure zero mutations or deletions to external production media directories during testing.
### Functional & Contract Compatibility
- [ ] All REST API endpoints (`/api/status`, `/api/files`, `/api/run`, `/api/rollback`, `/api/library`, `/api/quarantine`, etc.) remain 100% backward-compatible in request signatures and response schemas.
- [ ] Single-page web dashboard remains fully functional, retaining Folder Explorer (with `.txt`/`.srt` exclusions), Library catalog views, quarantine queue, all 30 CSS themes, RGB Chroma mode, and classification modals.
## Follow-up — 2026-09-07T16:36:03Z
Implement download client webhook integrations (qBittorrent, Transmission, SABnzbd, JDownloader, generic JSON), real-time filesystem directory monitoring, and an interactive saved show match-and-confirm system for Media Sorter.
Working directory: /md0/media-sorter
Integrity mode: development
## Requirements
### R1. Download Client Webhook Integrations
Provide dedicated HTTP webhook endpoints and handlers for popular download managers:
- qBittorrent (completion script / webhook payload parameters)
- Transmission (torrent-done script payload)
- SABnzbd (post-processing notification)
- JDownloader (Event Scripter webhook trigger)
- Generic JSON webhook (`POST /api/webhooks/download-complete` accepting path, title, category)
Each handler must parse incoming client payload data, validate that the target media path exists within configured source paths, and enqueue or trigger sorting for the completed item.
### R2. Real-Time Directory Filesystem Watcher
Implement a resilient background filesystem watcher attached to configured download directories:
- Monitor for file creation and move/rename completion events for supported media extensions.
- Include a configurable settling/debounce window so active or in-flight downloads are not processed prematurely.
- Gracefully start and stop with the FastAPI application lifespan, ignoring destination library folders, hidden temporary files, and partially downloaded files (`.part`, `.crdownload`, `!qB`).
### R3. Saved Shows & Match-to-Show Confirmation System
Implement a "Saved Shows" tracking system and smart match prompt:
- Allow users to explicitly mark, bookmark, or save shows in the library/database.
- When new incoming media is scanned (via watcher, webhook, or manual scan) and matches an already saved show, present a clear match option/prompt in the Web UI folder explorer rather than automatically moving without confirmation.
- Allow 1-click confirmation to route the matched files directly into the saved show's designated folder structure with resolved season and episode numbering.
### R4. Configuration & Web UI Management
Expose settings and controls in the Web UI and configuration file:
- Enable/disable filesystem watcher and configure debounce interval.
- Webhook documentation/URL generator with client-specific setup instructions and tokens if security is enabled.
- UI badges and confirmation banners in the Folder Explorer when incoming files match a saved show.
## Acceptance Criteria
### Webhook Handlers
- [ ] Automated tests verify each client endpoint (qBittorrent, Transmission, SABnzbd, JDownloader, generic JSON) parses valid inputs and dispatches the sorting pipeline.
- [ ] Invalid payloads, missing files, or out-of-boundary paths return appropriate HTTP error statuses (400, 404, or 422) without crashing the server.
### Filesystem Watcher
- [ ] Watcher starts and stops cleanly with server lifespan without orphaned background threads or memory leaks.
- [ ] Automated test proves writing a media file with a simulated debounce delay triggers processing only after file write completion.
- [ ] Temporary download extension files (`.part`, `!qB`) are ignored until finalized.
### Saved Shows & Match Confirmation
- [ ] Database model and API endpoints support saving/unsaving shows (`POST/DELETE /api/library/saved-shows` or equivalent).
- [ ] Incoming files matching a saved show trigger a distinct review/match prompt in the inspection response and Web UI.
- [ ] Confirming the match moves files into the existing show folder with correct season and episode structure.
### Test Suite & Regressions
- [ ] All existing 180 unit, integration, and benchmark tests continue to pass with 0 regressions.
- [ ] New unit and integration test coverage added for webhooks, filesystem watcher, and saved show matching.
+75
View File
@@ -0,0 +1,75 @@
# Project: Media Sorter Hardening & Test Expansion
## Architecture
The Media Sorter is a Python application (FastAPI + SQLAlchemy + SQLite) for automated and manual organization of TV shows, movies, and anime files.
- **Core Engine**: `media_sorter.sorter.MediaSorter`, `media_sorter.scanner.MediaScanner`, `media_sorter.tokenizer.FilenameTokenizer`, `media_sorter.classifier.MediaClassifier`, `media_sorter.namer.MediaNamer`, `media_sorter.executor.FileExecutor`.
- **Database Layer**: SQLite with WAL mode, SQLAlchemy ORM models (`BatchRecord`, `FileRecord`, `QuarantineRecord`, `ConfigAudit`, `OperationLog`, `LibraryItem`), Alembic migrations.
- **Service & Routing Layer**: FastAPI server orchestrating REST API endpoints, background auto-sorting workers, file inspection/clustering, and single-page dashboard.
- **Safety & Isolation**: Filesystem isolation traps, mock fixtures, non-mutating read endpoints, atomic move/rollback operations.
## Feature Inventory
| # | Feature | Description | Milestone | Source |
|---|---------|-------------|-----------|--------|
| F1 | Server Architecture Decoupling & UI Extraction | Extract 3,404-line HTML/CSS/JS dashboard into external templates/UI module, partition 27 endpoints into modular FastAPI router controllers (`routes/`), decouple file inspection into `file_service.py`, preserve backward-compatible facade re-exports in `server.py` | M3 | R1, survey |
| F2 | Web Dashboard & Theme Preservation | Retain full single-page web dashboard functionality: 5 tabs (Dashboard, Folder Explorer with `.txt`/`.srt` exclusions, Library catalog, Quarantine queue, Settings), 14 CSS themes with dropdown & swatches, 2 interactive modals | M3 | R1, survey |
| F3 | REST API Contract & Backward Compatibility | Maintain 100% backward compatibility for all 27 REST endpoints (`/api/status`, `/api/files`, `/api/run`, `/api/rollback`, `/api/library`, `/api/quarantine`, etc.) in request parameters, payloads, and response JSON schemas | M3 | R1, survey |
| F4 | Database Session Lifecycle & Engine Management | Eliminate broken `scoped_session` re-instantiation in `db.py`, provide proper connection pooling, clean session lifecycle scoping (`session.close()`, `remove()`), and thread-safe session factories | M1 | R2, survey |
| F5 | Database Transaction Safety & Concurrency Hardening | Implement concurrency locking on rollback and manual sort operations to prevent race conditions against active background sort runs; fix partial rollback batch state in `executor.py`; fix non-transactional `undo_item()` in `quarantine.py`; offload synchronous `sorter.run()` from async event loop in `auto_sort_worker` | M1 | R2, survey |
| F6 | Schema Synchronization & Library Catalog Integrity | Add missing `library_items` table to Alembic migrations; fix silent `AttributeError` on `op.status` in `sorter.py:271` so live sorts update `LibraryItem` entries; remove read-endpoint mutation in `GET /api/files` that cumulatively inflates `item_count` | M1 | R2, survey |
| F7 | Tokenizer Edge-Case Expansion | Enhance `FilenameTokenizer` regexes and parsing to handle Roman numerals (`Season II Episode IV`), ambiguous years (`1917`, `2001`, `2049`), anime titles with parentheses (`Fairy Tail (2014)`), and multi-part episode formats (`01-02`, `1x01-02`, `S01E01E02E03`) | M2 | R3, R4, survey |
| F8 | Error Handling, Input Validation & Security Guardrails | Sanitize user-provided filename strings in `/api/files/manual-sort`; validate storage boundaries in `/api/poster/local` to prevent arbitrary file reading; implement defensive structured error responses across all endpoints | M2 | R3, survey |
| F9 | Companion File Handling, Pairing, Exclusions & Cleanup | Verify and harden sidecar pairing (`.srt`, `.ass`, `.nfo`, etc.), language suffix preservation (`.en.srt`), Folder Explorer `.txt`/`.srt` exclusions, and directory cleanup routines | M2 | R1, R3, survey |
| F10 | Filesystem Safety Guardrails & Environment Isolation | Add root `tests/conftest.py` with autouse safety trap protecting production directories (`/md0/jdownloads`, `/md0/movies1`, `/md0/tv1`); decouple `Settings` from process-wide `os.environ` mutation in `POST /api/settings` to prevent cross-test contamination | M1 | R5, survey |
## Milestones
| # | Name | Scope | Dependencies | Status |
|---|------|-------|-------------|--------|
| E2E | E2E Testing Track | Requirement-driven opaque-box test suite (Tiers 1-4, ≥115 test cases), test runner, and `TEST_READY.md` publishing | None | PLANNED |
| M1 | Filesystem Safety Guardrails & Database Reliability Hardening | F4 (Session lifecycle), F5 (Transaction safety & concurrency), F6 (Schema sync & catalog updates), F10 (Filesystem safety trap in `tests/conftest.py` & env isolation) | None | PLANNED |
| M2 | Tokenizer Edge-Case Expansion & Input Validation Hardening | F7 (Roman numerals, ambiguous years, anime parentheses, multi-episodes), F8 (Input sanitization & security), F9 (Companion file pairing & exclusions) | M1 | PLANNED |
| M3 | Server Architectural Decoupling & UI Template Extraction | F1 (Monolith decomposition into modular routers and file service), F2 (Dashboard UI & theme preservation), F3 (REST API contract backward compatibility) | M1, M2 | PLANNED |
| M4 | Final Milestone: Full E2E Verification & Adversarial Coverage Hardening | Phase 1: Pass 100% of E2E test suite (Tiers 1-4) + 75 baseline tests. Phase 2: Adversarial Coverage Hardening (Tier 5) with Challenger loop to verify ≥90% coverage on core modules | E2E, M1, M2, M3 | PLANNED |
## Code Layout
- `src/media_sorter/`
- `config.py`: Configuration models and settings
- `models.py`: SQLAlchemy ORM models
- `db.py`: Database engine and session factory
- `scanner.py`: File scanning and sidecar pairing
- `tokenizer.py`: Filename tokenization and regex parsing
- `classifier.py`: Media categorization heuristics
- `namer.py`: Standardized filename generation
- `executor.py`: File moving, copying, and rollback execution
- `sorter.py`: High-level sorting orchestrator
- `library.py`: Media library catalog management
- `quarantine.py`: Quarantine management
- `file_service.py`: Decoupled file inspection and clustering service (extracted from server.py)
- `templates/`: Extracted single-page web dashboard HTML/CSS/JS template
- `routes/`: Modular FastAPI route controllers
- `status.py`, `batches.py`, `files.py`, `quarantine.py`, `library.py`, `settings.py`, `posters.py`
- `server.py`: Backward-compatible facade re-exporting `create_app` and core helpers
- `cli.py`: CLI command-line interface
- `tests/`
- `conftest.py`: Root safety guardrails, env sanitization, and filesystem protection traps
- `unit/`: Unit tests for core components
- `integration/`: Integration tests for server and pipeline workflows
- `e2e/`: Opaque-box requirement-driven E2E test suite (Tiers 1-4)
- `alembic/`: Database migrations
## Interface Contracts
### `server.py` Facade ↔ External Callers & Tests
- `create_app(settings: Optional[Settings] = None, engine: Optional[Engine] = None) -> FastAPI`
- `inspect_downloads_folder(downloads_dir: Path, movies_dir: Path, shows_dir: Path, ...) -> Dict[str, Any]`
- `list_files_in_dir(path: Path) -> List[Dict[str, Any]]`
- `cluster_unsure_files(files: List[Dict[str, Any]]) -> List[Dict[str, Any]]`
- `clean_detected_show_name(name: str) -> str`
- Request models: `RunRequest`, `RollbackRequest`, `ResolveRequest`, `BulkResolveRequest`, `BulkUndoRequest`, `ManualSortRequest`, `SortShowRequest`, `SortGroupRequest`, `SettingsUpdateRequest`.
### `db.py` ↔ Sorter & Routes
- `get_engine(database_url: str) -> Engine`
- `get_session_factory(engine: Engine) -> sessionmaker[Session]`
- `get_db_session(engine: Engine) -> Generator[Session, None, None]` (proper context manager with clean commit/rollback/close)
### `tokenizer.py` ↔ Scanner & Classifier
- `FilenameTokenizer.tokenize(path: Path) -> TokenizedMedia`
- Supports Roman numerals (`I..XX`), ambiguous years without title loss, anime titles with parentheses, and multi-part episodes (`multi_episodes: List[int]`).
+321
View File
@@ -0,0 +1,321 @@
# Media Sorter
A reliable, high-performance media classification and organization engine engineered with defensive data safety, atomic operations, dry-run simulation, operation journaling, transactional rollback, and review quarantine.
Supported media types:
- **Movies** (Feature films, scene releases, REMUX, UHD/4K, multi-part)
- **TV Shows** (Episodic series, multi-episode files, season packages)
- **Anime** (Fansub groups, absolute numbering, season/episode mapping)
- **Music** (Multi-disc albums, flac/mp3, ID3v2/Vorbis tags, track/artist tokens)
- **Audiobooks** (M4B, chapter tags, narrator metadata, multi-part)
- **Podcasts** (Dated releases, show prefixes, episode titles)
- **Documentaries** (Documentary flags, broadcast tags)
- **Home Videos & Photos** (EXIF datetime, camera models, smartphone naming schemes)
- **Sidecars & Companions** (Subtitles `.srt/.ass`, Artwork `poster/cover`, Metadata `.nfo`, Extras `-trailer/-sample`)
- **Archives & Unknowns** (Zip, rar, 7z, and unclassified files isolated safely)
---
## Key Safety Guarantees
1. **Zero Silent Data Loss**: Files are never deleted by default.
2. **Dry-Run by Default**: Operations always default to dry-run preview unless explicitly launched with `--live` or configured otherwise.
3. **No Guessing / Quarantine Queue**: If classification confidence falls below the configured threshold (default `0.75`), the file is placed into the Quarantine queue for human inspection.
4. **Collision & Conflict Prevention**: Destination paths are validated beforehand. If a collision is detected, the configurable policy (`rename_unique`, `replace_if_higher_quality`, `quarantine`, `skip`, or `error`) is triggered safely.
5. **Atomic Moves**: Moves on the same filesystem use `os.replace`. Moves across filesystems write to hidden temporary files (`.tmp_media_sorter_*`), verify integrity/size, atomically replace into destination, and only then remove source files.
6. **Transactional Journaling & Instant Rollback**: All operations are recorded in a SQLite WAL database (`OperationStatus.PLANNED` -> `IN_PROGRESS` -> `COMMITTED`). Any batch can be cleanly inverted with `media-sorter rollback --batch-id <id>`.
7. **Active Download / Lock Protection**: Automatically ignores files modified within the minimum file age (default 300s) or locked by downloading torrent/browser clients.
---
## Architecture
```
+-----------------------+
| Source Directories |
+-----------+-----------+
|
v
+-----------------------+
| Scanner (Lock/Age) |
+-----------+-----------+
|
+------------------------+------------------------+
| | |
v v v
+--------------------+ +-------------------+ +-------------------+
| Filename Tokenizer | | Media Analyzer | | External Provider |
| (Regex / Patterns) | | (Headers/Atoms) | | (TMDB/MusicBrainz)|
+----------+---------+ +---------+---------+ +---------+---------+
| | |
+------------------------+------------------------+
|
v
+-----------------------+
| Multi-Signal |
| Classifier & Scorer |
+-----------+-----------+
|
+---------------------+---------------------+
| (Confidence >= 0.75)| | (Confidence < 0.75)
v v
+-----------------------+ +-----------------------+
| Namer & Path Sanitizer| | Quarantine Manager |
+-----------+-----------+ +-----------------------+
|
v
+-----------------------+
| Executor & Journal | <==== SQLite WAL Database (media_sorter.db)
+-----------+-----------+
|
+-----------+-----------+
| Organized Destination |
+-----------------------+
```
---
## Quickstart
### 1. Installation
```bash
# Clone repository
git clone https://github.com/example/media-sorter.git
cd media-sorter
# Create virtual environment and install
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
```
### 2. Configuration (.env)
Media Sorter supports simple, editable settings directly in a `.env` file in the project root:
```ini
# .env
DOWNLOADS_DIR=./downloads # Source folder where downloads arrive
MOVIES_DIR=./movies # Destination for Movies
SHOWS_DIR=./shows # Destination for TV Series
DRY_RUN=false # false = move files live; true = simulation preview
ACTION=move # move | copy | link | hardlink
CONFIDENCE_THRESHOLD=0.75 # Minimum classification confidence (0.0 - 1.0)
MIN_FILE_AGE_SECONDS=0 # Ignore files modified within N seconds
SERVER_HOST=0.0.0.0
SERVER_PORT=8085
DATABASE_PATH=media_sorter.db
```
Check your active configuration at any time:
```bash
media-sorter config show
```
### 3. Web Dashboard & Management UI
Start the responsive, modern Web Management Dashboard:
```bash
media-sorter server
```
**Using the Antigravity Bot in Direct Messages**
You can also interact with the Media Sorter assistant bot via direct messages (DMs). Send commands such as `scan`, `organize`, `rollback`, or `config show` directly to the bot, and it will reply with results, status updates, or interactive prompts. This provides a convenient way to manage your media library without opening the web UI.
**Bot Owner Configuration**
Set the Discord bot owner ID in `config.example.yaml` (or your own config file) under the `owner` field. The value should be the numeric Discord user ID of the bot’s owner. The bot will treat this user as having owner‑only privileges.
Open `http://localhost:8085` (or `http://<your-host-ip>:8085`) in your browser. The web UI includes:
- **Dashboard & Activity**: Live metrics, mode badge, 1-click Dry-Run Preview & Live Sort buttons, 1-click Server Restart button, and past batch history.
- **Folder Explorer**: Live view of files in `downloads/`, `movies/`, and `shows/` with file sizes and timestamps.
- **Quarantine Review**: Visual approval queue for ambiguous media with 1-click "Approve as Movie" or "Approve as Show".
- **Settings & .env**: Interactive settings form where you can update directory paths, toggle dry-run mode, and save directly to `.env`.
### 4. Running with PM2 (Production Process Manager)
Media Sorter includes an [`ecosystem.config.js`](file:///md0/media-sorter/ecosystem.config.js) file for daemonizing under PM2:
```bash
# Start Media Sorter with PM2
pm2 start ecosystem.config.js
# View status
pm2 status
# View live stream logs
pm2 logs media-sorter
# Restart or reload
pm2 restart media-sorter
# Enable PM2 to auto-start on machine boot
pm2 save
pm2 startup
```
The web dashboard's **🔄 Restart Server** button connects natively with PM2: clicking it restarts the process and automatically refreshes the web page once back online.
### 5. Scan & Preview (Dry-Run)
Preview classification without touching any files:
```bash
media-sorter scan
```
Generate a dry-run batch preview:
```bash
media-sorter organize --dry-run
```
### 5. Execute Live Organization
Sort files from `downloads/` into `movies/` or `shows/`:
```bash
media-sorter organize --live
```
### 6. Instant Rollback
If you ever need to undo an organization run:
```bash
# Revert latest batch
media-sorter rollback
# Revert specific batch
media-sorter rollback --batch-id <batch-uuid>
```
### 7. Review Quarantine Queue
List and resolve files requiring manual verification via CLI:
```bash
media-sorter quarantine list
media-sorter quarantine resolve 1 --category movie
```
---
## CLI Reference
| Command | Description |
|---|---|
| `media-sorter scan` | Discover and analyze source media, displaying classification table |
| `media-sorter organize` | Execute organization or dry-run preview (`--dry-run` or `--live`) |
| `media-sorter rollback` | Roll back a batch and restore files to source paths |
| `media-sorter history` | View audit trail of past batches and metrics |
| `media-sorter quarantine list` | List items pending manual human review |
| `media-sorter quarantine resolve` | Approve or reclassify a quarantined item |
| `media-sorter server` | Start FastAPI REST API and web management dashboard |
| `media-sorter config show` | Display active configuration settings in YAML |
| `media-sorter config init` | Generate a starter configuration file |
---
## Database & State Tracking
The system utilizes SQLite in **Write-Ahead Logging (WAL)** mode with `PRAGMA synchronous=NORMAL` and `PRAGMA foreign_keys=ON`:
- `batches`: High-level run records with status (`IN_PROGRESS`, `COMPLETED`, `ROLLED_BACK`), dry-run indicator, counts of moved, skipped, failed, and quarantined files.
- `operations`: Fine-grained journal entries tracking `src`, `dst`, `action` (`move`, `copy`, `link`, `hardlink`), `src_hash`, `dst_hash`, `backup_path`, diagnostic details, and timestamps.
- `files`: File fingerprint cache (`path`, `size`, `mtime`, `hash`) to avoid redundant metadata probing on unchanged files.
- `quarantine`: Audit log for low-confidence or conflicting files holding reason, signals, and resolution states.
### Migrations with Alembic
Run database migrations:
```bash
alembic upgrade head
```
Create a new migration:
```bash
alembic revision --autogenerate -m "Add custom column"
```
---
## Deployment
### Docker
Build and run with Docker Compose:
```bash
docker-compose up -d
```
Check health status:
```bash
curl -f http://localhost:8080/api/status
```
### Systemd Service
1. Copy repository to `/opt/media-sorter`.
2. Copy `media-sorter.service` to `/etc/systemd/system/media-sorter.service`.
3. Enable and start:
```bash
sudo systemctl daemon-reload
sudo systemctl enable --now media-sorter
sudo systemctl status media-sorter
```
---
## Backup & Upgrade Instructions
### Database Backup
Because SQLite uses WAL mode, use the standard SQLite online backup or VACUUM INTO command:
```bash
# Safe hot-backup of live database
sqlite3 media_sorter.db ".backup 'media_sorter.backup.db'"
```
Or backup directory before upgrading:
```bash
cp media_sorter.db media_sorter.db.bak
```
### Upgrading
1. Pull latest release:
```bash
git pull origin main
```
2. Update dependencies:
```bash
pip install -e .
```
3. Run Alembic schema migrations:
```bash
alembic upgrade head
```
4. Restart service:
```bash
sudo systemctl restart media-sorter
```
---
## Testing
Run the full test suite (unit tests, integration tests, and Hypothesis property-based fuzz tests):
```bash
pytest -v
```
+226
View File
@@ -0,0 +1,226 @@
# Media Sorter E2E Benchmark Test Suite & Offline Runner (TEST_READY)
**Milestone**: E2E Benchmark Suite (Requirement R2 / Feature F13)
**Date**: 2026-09-05
**Working Directory**: `/md0/media-sorter`
**Test Location**: `tests/benchmark/`
**Status**: COMPLETE & VERIFIED
---
## 1. Executive Summary
Requirement R2 requires developing a repeatable, isolated automated verification suite testing filename tokenization, media classification, and destination path resolution against a diverse benchmark dataset of real-world media filename patterns.
The E2E Benchmark Test Suite and Offline Diagnostic Runner have been established:
- **Test Inventory**: 64 real-world media filename patterns spanning 6 distinct domains.
- **Pure In-Memory Offline Execution**: Operates with 0 external network requests, 0 filesystem mutations, and zero risk to host media storage (`/md0/jdownloads`, `/md0/movies1`, `/md0/tv1`).
- **Execution Performance**: Full 64-case suite executes in ~107 milliseconds.
- **Diagnostic Transparency**: Field-by-field diff reporting across category, title, year, season, episode, multi-episodes, date, edition, part, group, and destination subpath.
- **Zero Regressions**: All 82 existing unit, integration, safety, and defect tests continue to pass with a 100% pass rate.
---
## 2. Test Runner Invocation Commands
### 2.1 Pytest Benchmark Runner
Executes the parametrized pytest test suite (`@pytest.mark.parametrize("case", BENCHMARK_CASES, ids=lambda c: c.id)`):
```bash
.venv/bin/pytest tests/benchmark/test_benchmark.py
```
To run specific domains or cases:
```bash
.venv/bin/pytest tests/benchmark/test_benchmark.py -k "TV"
.venv/bin/pytest tests/benchmark/test_benchmark.py -k "ANIME"
.venv/bin/pytest tests/benchmark/test_benchmark.py -k "MOVIE"
```
### 2.2 Standalone CLI Benchmark Runner
Executes the offline benchmark runner with Rich terminal formatting and JSON export:
```bash
.venv/bin/python -m tests.benchmark.runner
```
Or directly:
```bash
.venv/bin/python tests/benchmark/runner.py
```
#### Supported CLI Options:
| Flag | Description | Example |
|---|---|---|
| `--domain <domain>` | Filter execution to a single domain | `.venv/bin/python -m tests.benchmark.runner --domain Anime` |
| `--case <case_id>` | Filter execution to a single benchmark case | `.venv/bin/python -m tests.benchmark.runner --case TV-01` |
| `--json-output <path>` | Custom path for exported summary JSON | `--json-output custom_summary.json` |
| `--no-diffs` | Print summary table only without failure diffs | `--no-diffs` |
| `--verbose`, `-v` | Display output for all cases (including passed) | `-v` |
| `--strict` | Exit with status code 1 if any case fails | `--strict` |
### 2.3 Existing Regression Test Suite
Verify that all unit and integration tests remain intact:
```bash
.venv/bin/pytest tests/unit/ tests/integration/
```
Result: **82 passed in ~1.67s** (100% pass rate).
---
## 3. Benchmark Domain Coverage Matrix (64 Cases)
The benchmark inventory defines authoritative ground truth for 64 real-world patterns across 6 domains:
### Domain 1: Standard TV (11 Cases)
| ID | Filename | Edge Case Type | Expected Category | Expected Season/Episode | Expected Destination Subpath | Baseline Status |
|---|---|---|---|---|---|---|
| `TV-01` | `Breaking.Bad.S05E14.Ozymandias.1080p.BluRay.x264-ROVERS.mkv` | Standard SxxExx | tv | S05E14 | `TV Shows/Breaking Bad/Season 05/Breaking Bad - S05E14.mkv` | Pending M3 format (`_` vs ` - `) |
| `TV-02` | `Stranger.Things.S04E01-E02.Chapter.One.720p.WEB-DL.mkv` | Multi-ep hyphen range | tv | S04E01-E02 | `TV Shows/Stranger Things/Season 04/Stranger Things - S04E01-E02.mkv` | Pending M2 multi-ep & M3 format |
| `TV-03` | `House.M.D.S03E01E02.1080p.mkv` | Concatenated multi-ep | tv | S03E01-E02 | `TV Shows/House M D/Season 03/House M D - S03E01-E02.mkv` | Pending M2 multi-ep & M3 format |
| `TV-04` | `The.Wire.1x09.HDTV.mkv` | Scene 1x09 | tv | S01E09 | `TV Shows/The Wire/Season 01/The Wire - S01E09.mkv` | Pending M3 format (`_` vs ` - `) |
| `TV-05` | `The.Office.2x01-02.mkv` | Scene multi-ep range | tv | S02E01-E02 | `TV Shows/The Office/Season 02/The Office - S02E01-E02.mkv` | Pending M2 multi-ep disambiguation |
| `TV-06` | `Rome.Season.II.Episode.IV.mkv` | Roman numerals | tv | S02E04 | `TV Shows/Rome/Season 02/Rome - S02E04.mkv` | Pending M2 Roman numeral parser |
| `TV-07` | `Doctor Who Season 5 Episode 1 Eleventh Hour.mkv` | Word season & episode | tv | S05E01 | `TV Shows/Doctor Who/Season 05/Doctor Who - S05E01.mkv` | Pending M3 format (`_` vs ` - `) |
| `TV-08` | `Succession.S02.Complete.1080p.WEB-DL.mkv` | Complete season pack | tv | S02 (Pack) | `TV Shows/Succession/Season 02/Succession - Season 02.mkv` | Pending M2 season pack tokenizer |
| `TV-09` | `Game.of.Thrones.S08E03.720p.HDTV.x264-AVS.mkv` | Scene release group | tv | S08E03 | `TV Shows/Game of Thrones/Season 08/Game of Thrones - S08E03.mkv` | Pending M3 format (`_` vs ` - `) |
| `TV-10` | `Chernobyl.S01E05.Vichnaya.Pamyat.1080p.mkv` | Episode title in stem | tv | S01E05 | `TV Shows/Chernobyl/Season 01/Chernobyl - S01E05.mkv` | Pending M3 format (`_` vs ` - `) |
| `TV-11` | `Friends.S06E15-E16.The.One.That.Could.Have.Been.mkv` | Double-digit multi-ep | tv | S06E15-E16 | `TV Shows/Friends/Season 06/Friends - S06E15-E16.mkv` | Pending M2 multi-ep & M3 format |
### Domain 2: Anime (11 Cases)
| ID | Filename | Edge Case Type | Expected Category | Expected Episode / Group | Expected Destination Subpath | Baseline Status |
|---|---|---|---|---|---|---|
| `ANIME-01` | `[SubsPlease] Frieren - Beyond Journey's End - 01 (1080p) [ABCD1234].mkv` | Fansub brackets & CRC | anime | Ep 01 / SubsPlease | `Anime/Frieren - Beyond Journey's End/Frieren - Beyond Journey's End - 01 [SubsPlease].mkv` | Pending M3 Anime destination routing |
| `ANIME-02` | `[SubsPlease] 葬送のフリーレン - 12 (1080p) [98E7B1A2].mkv` | Unicode / Kanji title | anime | Ep 12 / SubsPlease | `Anime/葬送のフリーレン/葬送のフリーレン - 12 [SubsPlease].mkv` | Pending M3 Anime destination routing |
| `ANIME-03` | `[HorribleSubs] Fairy Tail (2014) - 176 [720p].mkv` | Parenthesized title year | anime | Ep 176 / HorribleSubs | `Anime/Fairy Tail (2014)/Fairy Tail (2014) - 176 [HorribleSubs].mkv` | Pending M2 Anime parenthesized title |
| `ANIME-04` | `[Erai-raws] One Piece - 1088 [1080p].mkv` | 4-digit absolute numbering | anime | Ep 1088 / Erai-raws | `Anime/One Piece/One Piece - 1088 [Erai-raws].mkv` | Pending M2 absolute number routing |
| `ANIME-05` | `[SubsPlease] Dungeon Meshi - 01-02 (1080p).mkv` | Anime multi-episode range | anime | Ep 01-02 / SubsPlease | `Anime/Dungeon Meshi/Dungeon Meshi - 01-02 [SubsPlease].mkv` | Pending M2 multi-ep anime |
| `ANIME-06` | `[TaigaSubs] Attack on Titan OVA - 01 [720p].mkv` | Anime OVA special | anime | Ep 01 / TaigaSubs | `Anime/Attack on Titan OVA/Attack on Titan OVA - 01 [TaigaSubs].mkv` | Pending M2 OVA parser |
| `ANIME-07` | `Naruto Episode 207 The Supposed Sealed Ability.mkv` | Standalone episode keyword | anime | Ep 207 | `Anime/Naruto/Naruto - 207.mkv` | Pending M2 anime keyword classification |
| `ANIME-08` | `[Judas] Fate Stay Night - Heaven's Feel - I. Presage Flower [BD 1080p].mkv` | Roman numeral sub-title | anime | Movie / Judas | `Anime/Fate Stay Night - Heaven's Feel - I. Presage Flower/...` | Pending M2 Roman numeral & anime routing |
| `ANIME-09` | `BLEACH - Sennen Kessen-hen - 27 [E89717B7].mkv` | No-group fansub with CRC | anime | Ep 27 | `Anime/BLEACH - Sennen Kessen-hen/BLEACH - Sennen Kessen-hen - 27.mkv` | Pending M2 fansub syntax & M3 routing |
| `ANIME-10` | `[Erai-raws] Jujutsu Kaisen 2nd Season - 14 [1080p][Multiple Subtitle].mkv` | Cour / season tags | anime | Ep 14 / Erai-raws | `Anime/Jujutsu Kaisen 2nd Season/Jujutsu Kaisen 2nd Season - 14 [Erai-raws].mkv` | Pending M2 cour pattern |
| `ANIME-11` | `[SubsPlease] Mushoku Tensei S2 - 18 (1080p) [F28B1452].mkv` | Short season notation S2 | anime | Ep 18 / SubsPlease | `Anime/Mushoku Tensei S2/Mushoku Tensei S2 - 18 [SubsPlease].mkv` | Pending M2 S2 anime tokenizer |
### Domain 3: Movies (14 Cases)
| ID | Filename | Edge Case Type | Expected Category | Expected Year / Edition / Part | Expected Destination Subpath | Baseline Status |
|---|---|---|---|---|---|---|
| `MOVIE-01` | `Inception.2010.1080p.BluRay.x264-FraMeSToR.mkv` | Standard release year | movie | 2010 | `Movies/Inception (2010)/Inception (2010).mkv` | Pending M3 movie template name |
| `MOVIE-02` | `1917.2019.1080p.BluRay.x264.mkv` | Numerical title with year | movie | 2019 (Title: 1917) | `Movies/1917 (2019)/1917 (2019).mkv` | Pending M2 year disambiguation |
| `MOVIE-03` | `2001.A.Space.Odyssey.1968.REMASTERED.1080p.mkv` | Numerical title & edition | movie | 1968 / Remastered | `Movies/2001 A Space Odyssey (1968)/... [Remastered].mkv` | Pending M2 numerical title & edition |
| `MOVIE-04` | `Blade.Runner.2049.2017.2160p.UHD.BluRay.x265.mkv` | Future year in title | movie | 2017 (Title: Blade Runner 2049) | `Movies/Blade Runner 2049 (2017)/...` | Pending M2 year disambiguation |
| `MOVIE-05` | `Wonder.Woman.1984.2020.1080p.WEB-DL.mkv` | Year in title | movie | 2020 (Title: Wonder Woman 1984) | `Movies/Wonder Woman 1984 (2020)/...` | Pending M2 year disambiguation |
| `MOVIE-06` | `Class.of.1999.1990.720p.mkv` | Past year in title | movie | 1990 (Title: Class of 1999) | `Movies/Class of 1999 (1990)/...` | Pending M2 year disambiguation |
| `MOVIE-07` | `Titanic.1997.DVD.CD1.avi` | Multi-part movie CD1 | movie | 1997 / Part 1 | `Movies/Titanic (1997)/Titanic (1997) [Pt.1].avi` | Pending M2 part parser & M3 path |
| `MOVIE-08` | `Titanic.1997.DVD.CD2.avi` | Multi-part movie CD2 | movie | 1997 / Part 2 | `Movies/Titanic (1997)/Titanic (1997) [Pt.2].avi` | Pending M2 part parser & M3 path |
| `MOVIE-09` | `Kill.Bill.Vol.1.2003.1080p.BluRay.mkv` | Volume number in title | movie | 2003 (Title: Kill Bill Vol 1) | `Movies/Kill Bill Vol 1 (2003)/...` | Pending M3 movie template name |
| `MOVIE-10` | `The.Lord.of.the.Rings.The.Fellowship.of.the.Ring.2001.Extended.1080p.mkv` | Extended edition | movie | 2001 / Extended | `Movies/The Lord of the Rings... [Extended].mkv` | Pending M2 edition parser |
| `MOVIE-11` | `Aliens.1986.Directors.Cut.1080p.BluRay.mkv` | Director's Cut edition | movie | 1986 / Director's Cut | `Movies/Aliens (1986)/Aliens (1986) [Director's Cut].mkv` | Pending M2 edition parser |
| `MOVIE-12` | `Gladiator.2000.Remastered.1080p.BluRay.mkv` | Remastered edition | movie | 2000 / Remastered | `Movies/Gladiator (2000)/Gladiator (2000) [Remastered].mkv` | Pending M2 edition parser |
| `MOVIE-13` | `Seven.Samurai.1954.Criterion.Collection.1080p.BluRay.mkv` | Criterion edition | movie | 1954 / Criterion | `Movies/Seven Samurai (1954)/Seven Samurai (1954) [Criterion].mkv` | Pending M2 edition parser |
| `MOVIE-14` | `Blade.Runner.1982.Final.Cut.2160p.UHD.mkv` | Final Cut edition | movie | 1982 / Final Cut | `Movies/Blade Runner (1982)/Blade Runner (1982) [Final Cut].mkv` | Pending M2 edition parser |
### Domain 4: Specials & Extras (9 Cases)
| ID | Filename | Edge Case Type | Expected Category | Expected Season/Episode | Expected Destination Subpath | Baseline Status |
|---|---|---|---|---|---|---|
| `SPECIAL-01` | `The.Office.S00E01.The.Outtakes.mkv` | Season 00 TV special | tv | S00E01 | `TV Shows/The Office/Season 00/The Office - S00E01.mkv` | Pending M3 Season 00 falsy bug fix |
| `SPECIAL-02` | `Doctor.Who.S00E25.The.Day.of.the.Doctor.1080p.mkv` | High episode Season 00 | tv | S00E25 | `TV Shows/Doctor Who/Season 00/Doctor Who - S00E25.mkv` | Pending M3 Season 00 falsy bug fix |
| `SPECIAL-03` | `Breaking.Bad.S05E00.Special.mkv` | Mid-season special E00 | tv | S05E00 | `TV Shows/Breaking Bad/Season 05/Breaking Bad - S05E00.mkv` | Pending M3 Season 00 falsy bug fix |
| `SPECIAL-04` | `Inception.2010-behindthescenes.mkv` | Behind the scenes extra | movie | 2010 | `Movies/Inception (2010)/Inception (2010)-behindthescenes.mkv` | Pending M3 extra suffix preservation |
| `SPECIAL-05` | `The.Matrix.1999-featurette.mkv` | Movie featurette extra | movie | 1999 | `Movies/The Matrix (1999)/The Matrix (1999)-featurette.mkv` | Pending M3 extra suffix preservation |
| `SPECIAL-06` | `Interstellar.2014-deleted.mkv` | Deleted scenes extra | movie | 2014 | `Movies/Interstellar (2014)/Interstellar (2014)-deleted.mkv` | Pending M3 extra suffix preservation |
| `SPECIAL-07` | `Interstellar.2014-trailer.mp4` | Movie trailer extra | movie | 2014 | `Movies/Interstellar (2014)/Interstellar (2014)-trailer.mp4` | Pending M3 extra suffix preservation |
| `SPECIAL-08` | `Game.of.Thrones.S00E02.A.Day.in.the.Life.mkv` | Season 00 with title | tv | S00E02 | `TV Shows/Game of Thrones/Season 00/Game of Thrones - S00E02.mkv` | Pending M3 Season 00 falsy bug fix |
| `SPECIAL-09` | `Sherlock.S00E01.Many.Happy.Returns.mkv` | TV special prequel | tv | S00E01 | `TV Shows/Sherlock/Season 00/Sherlock - S00E01.mkv` | Pending M3 Season 00 falsy bug fix |
### Domain 5: Daily / Dated Shows (9 Cases)
| ID | Filename | Edge Case Type | Expected Category | Expected Date | Expected Destination Subpath | Baseline Status |
|---|---|---|---|---|---|---|
| `DAILY-01` | `The.Daily.Show.2024-01-15.1080p.HDTV.mkv` | ISO dated TV | tv | 2024-01-15 | `TV Shows/The Daily Show/Season 2024/The Daily Show - 2024-01-15.mkv` | Pending M2 daily regex & M3 TV routing |
| `DAILY-02` | `The.Tonight.Show.Starring.Jimmy.Fallon.2024.03.12.720p.mkv` | Dot-dated TV | tv | 2024-03-12 | `TV Shows/The Tonight Show.../Season 2024/... - 2024-03-12.mkv` | Pending M2 daily regex & M3 TV routing |
| `DAILY-03` | `Last.Week.Tonight.with.John.Oliver.2023-11-05.1080p.mkv` | Weekly dated show | tv | 2023-11-05 | `TV Shows/Last Week Tonight.../Season 2023/... - 2023-11-05.mkv` | Pending M2 daily regex & M3 TV routing |
| `DAILY-04` | `Late.Night.with.Seth.Meyers.2024_02_20.720p.mkv` | Underscore dated TV | tv | 2024-02-20 | `TV Shows/Late Night.../Season 2024/... - 2024-02-20.mkv` | Pending M2 daily regex & M3 TV routing |
| `DAILY-05` | `The Daily - 2026-03-12 - The Sunday Read.mp3` | Dated podcast | podcast | 2026-03-12 | `Podcasts/The Daily/2026/The Daily - 2026-03-12 - The Sunday Read.mp3` | Pending M2 podcast date & classifier |
| `DAILY-06` | `Jimmy.Kimmel.Live.2024-04-18.720p.HDTV.mkv` | Daily show ISO | tv | 2024-04-18 | `TV Shows/Jimmy Kimmel Live/Season 2024/Jimmy Kimmel Live - 2024-04-18.mkv` | Pending M2 daily regex & M3 TV routing |
| `DAILY-07` | `PBS.NewsHour.2024.05.01.720p.mkv` | News dot date | tv | 2024-05-01 | `TV Shows/PBS NewsHour/Season 2024/PBS NewsHour - 2024-05-01.mkv` | Pending M2 daily regex & M3 TV routing |
| `DAILY-08` | `The.Late.Show.with.Stephen.Colbert.2024-02-14.1080p.mkv` | Late night show ISO | tv | 2024-02-14 | `TV Shows/The Late Show.../Season 2024/... - 2024-02-14.mkv` | Pending M2 daily regex & M3 TV routing |
| `DAILY-09` | `NPR.News.Now.2024-06-10.mp3` | Podcast daily news | podcast | 2024-06-10 | `Podcasts/NPR News Now/2024/NPR News Now - 2024-06-10.mp3` | Pending M2 podcast date & classifier |
### Domain 6: Messy & Complex (10 Cases)
| ID | Filename | Edge Case Type | Expected Category | Expected Attributes | Expected Destination Subpath | Baseline Status |
|---|---|---|---|---|---|---|
| `MESSY-01` | `Amélie.2001.PROPER.REMASTERED.1080p.BluRay.x264-CiNEFiLE.mkv` | Accented characters | movie | 2001 / Remastered / CiNEFiLE | `Movies/Amélie (2001)/Amélie (2001) [Remastered].mkv` | Pending M2 edition & M3 template |
| `MESSY-02` | `Wolfs.2024.1080p.Apple.TV.WEB-DL.DDP5.1.Atmos.H.264.mkv` | Apple TV scene tag | movie | 2024 | `Movies/Wolfs (2024)/Wolfs (2024).mkv` | Pending M2 Apple.TV spec stripping |
| `MESSY-03` | `Gladiator.II.2024.1080p.HDTV.x264-[rartv].mkv` | Roman numeral title | movie | 2024 / rartv | `Movies/Gladiator II (2024)/Gladiator II (2024).mkv` | Pending M2 bracket group parsing |
| `MESSY-04` | `Interstellar.1920x1080.mkv` | Resolution dimensions | movie | Title: Interstellar | `Movies/Interstellar/Interstellar.mkv` | Pending M2 dimension stripping & quarantine |
| `MESSY-05` | `[YTS.MX] Movie Title - 2024 [1080p].mkv` | Bracket movie group | movie | 2024 / YTS.MX | `Movies/Movie Title (2024)/Movie Title (2024).mkv` | Pending M2 bracket group non-anime |
| `MESSY-06` | `The.Dark.Knight.2008.1080p.forced.srt` | Subtitle forced tag | subtitle | 2008 / Subtitle | `Movies/The Dark Knight (2008)/The Dark Knight (2008).forced.srt` | **PASSED** (100% match) |
| `MESSY-07` | `Show: "Special" <Episode> \| 1?.mkv` | Forbidden characters | tv | Ep 1 / Sanitized | `TV Shows/Show Special Episode 1/Season 01/... - S01E01.mkv` | Pending M2 character sanitization |
| `MESSY-08` | `CON.mp4` | Windows reserved device | home_video | CON -> _CON | `Home Videos/2026/2026-01 - Event/_CON.mp4` | Pending M2 reserved name handling |
| `MESSY-09` | ` Messy Show . S01E01 . 1080p .mkv` | Whitespace & dot padding | tv | S01E01 | `TV Shows/Messy Show/Season 01/Messy Show - S01E01.mkv` | Pending M2 whitespace collapse |
| `MESSY-10` | `Show_Name__2022__S02E03__HDTV.mkv` | Consecutive underscores | tv | 2022 / S02E03 | `TV Shows/Show Name/Season 02/Show Name - S02E03.mkv` | Pending M2 underscore normalization |
---
## 4. Initial Pass/Fail Baseline Report
### 4.1 Empirical Baseline Summary Table
Measured against codebase at commit baseline (`M0 / Phase 0`) prior to M2/M3 implementation:
| Domain | Total Cases | Passed | Failed | Pass Rate (%) | Avg Duration |
|---|---|---|---|---|---|
| **Standard TV** | 11 | 0 | 11 | 0.0% | 1.75 ms |
| **Anime** | 11 | 0 | 11 | 0.0% | 1.66 ms |
| **Movies** | 14 | 0 | 14 | 0.0% | 1.66 ms |
| **Specials & Extras** | 9 | 0 | 9 | 0.0% | 1.65 ms |
| **Daily / Dated Shows** | 9 | 0 | 9 | 0.0% | 1.65 ms |
| **Messy & Complex** | 10 | 1 | 9 | 10.0% | 1.65 ms |
| **OVERALL** | **64** | **1** | **63** | **1.6%** | **106.9 ms total** |
### 4.2 Gap Classification & Milestone Resolution Roadmap
The 63 baseline test failures cleanly delineate the exact functional requirements scheduled for Milestones M2 and M3:
#### Gap Group 1: Numerical Title & Year Ambiguity (M2: Feature F3)
- **Observed Defect**: `RE_YEAR` searches left-to-right from beginning of filename stem. For filenames such as `1917.2019`, `2001.A.Space.Odyssey.1968`, `Blade.Runner.2049.2017`, `Wonder.Woman.1984.2020`, the first 4-digit number is falsely extracted as the release year, wiping out or corrupting the movie title.
- **Affected Benchmark Cases**: `MOVIE-02`, `MOVIE-03`, `MOVIE-04`, `MOVIE-05`, `MOVIE-06`.
- **Target Resolution**: Milestone M2 implements right-to-left year detection (`RE_SCENE_YEAR`) and parenthesis year matching (`RE_PAREN_YEAR`).
#### Gap Group 2: Movie Editions & Multi-Part Split Files (M2: Feature F4)
- **Observed Defect**: `TokenizedFilename` lacks `edition` and `part` fields. Tokens such as `Extended`, `Director's Cut`, `Remastered`, `Criterion`, `Final Cut` are ignored or pollute the movie title. Multi-part files (`CD1`, `CD2`) produce identical destination paths, causing overwrite collisions.
- **Affected Benchmark Cases**: `MOVIE-03`, `MOVIE-07`, `MOVIE-08`, `MOVIE-10`, `MOVIE-11`, `MOVIE-12`, `MOVIE-13`, `MOVIE-14`, `MESSY-01`.
- **Target Resolution**: Milestone M2 adds `RE_EDITION` and `RE_MOVIE_PART`, populating `edition`, `part`, `part_label` in `TokenizedFilename`.
#### Gap Group 3: TV Roman Numerals, Multi-Episodes & Season Packs (M2: Feature F5)
- **Observed Defect**: `RE_SEASON_EPISODE` does not parse Roman numerals (`Season II Episode IV`), multi-episode ranges (`S04E01-E02`, `2x01-02`, `S03E01E02`), or season packs (`Succession.S02.Complete`). These fall back to movie classification or quarantine.
- **Affected Benchmark Cases**: `TV-02`, `TV-03`, `TV-05`, `TV-06`, `TV-08`, `TV-11`.
- **Target Resolution**: Milestone M2 upgrades `RE_SEASON_EPISODE` to support Roman numerals, multi-episode ranges, and season packs.
#### Gap Group 4: Daily / Dated Broadcast TV Shows (M2: Feature F6)
- **Observed Defect**: Shows with broadcast air dates (`2024-01-15`, `2024.03.12`, `2024_02_20`) contain year stamps but no episode numbers. `classifier.py` awards movie points for the year and misclassifies daily TV shows as movies, causing daily broadcasts to overwrite each other.
- **Affected Benchmark Cases**: `DAILY-01`, `DAILY-02`, `DAILY-03`, `DAILY-04`, `DAILY-06`, `DAILY-07`, `DAILY-08`.
- **Target Resolution**: Milestone M2 adds `RE_DAILY_DATE` to populate `is_daily` and `air_date`, steering `classifier.py` to classify dated broadcasts as `tv`.
#### Gap Group 5: Anime Absolute Numbering & Parenthesized Titles (M2: Feature F7)
- **Observed Defect**: `RE_ANIME_RELEASE` rejects parenthesized titles (`Fairy Tail (2014)`) and fails to handle 4-digit episode numbers (`One Piece - 1088`) or OVA releases.
- **Affected Benchmark Cases**: `ANIME-03`, `ANIME-04`, `ANIME-05`, `ANIME-06`, `ANIME-07`, `ANIME-08`, `ANIME-10`, `ANIME-11`.
- **Target Resolution**: Milestone M2 upgrades `RE_ANIME_RELEASE` to support parenthesized titles, cour numbers, and absolute numbering.
#### Gap Group 6: Season 00 Specials Falsy Bug & Template Formatting (M3: Feature F9)
- **Observed Defect**: In `namer.py:164`, `(tokens.season or 1)` evaluates `0` as falsy, converting Season 0 specials into Season 1 and overwriting pilots (`S01E01`). In `config.py:121`, the TV template uses underscore `show_name_{season_episode}` instead of the standard hyphen separator `show_name - {season_episode}`.
- **Affected Benchmark Cases**: `TV-01`, `TV-04`, `TV-07`, `TV-09`, `TV-10`, `SPECIAL-01`, `SPECIAL-02`, `SPECIAL-03`, `SPECIAL-08`, `SPECIAL-09`.
- **Target Resolution**: Milestone M3 fixes `tokens.season is not None` in `namer.py` and standardizes TV destination templates.
#### Gap Group 7: Anime Destination Routing & Library Indexing (M3: Feature F11)
- **Observed Defect**: Anime destination formatting routes files to `Season 01/One Piece_S01E1085 [UnknownGroup].mkv` rather than absolute episode format. In `sorter.py:274`, `dst_p.relative_to(shows_dir)` raises `ValueError` on anime paths, silently dropping anime items from library sync.
- **Affected Benchmark Cases**: `ANIME-01`, `ANIME-02`, `ANIME-04`, `ANIME-09`.
- **Target Resolution**: Milestone M3 adds absolute episode template rendering, stops injecting `[UnknownGroup]`, and fixes relative anime path indexing.
---
## 5. Filesystem Safety & Test Independence Verification
1. **Pure In-Memory Guarantee**:
All benchmark executions in `test_benchmark.py` and `runner.py` operate on pure in-memory `ScannedFile`, `MediaMetadata`, and `TokenizedFilename` objects. Zero files are opened for write/append, and zero temporary directories are created on disk.
2. **Path Trapping Compliance**:
All paths passed to `MediaNamer` use isolated mock base directories (`/test_dest`). Host media libraries (`/md0/jdownloads`, `/md0/movies1`, `/md0/tv1`) are never accessed or mutated.
3. **Deterministic Output**:
Every case execution produces deterministic tokenization and classification results across repeated invocations.
4. **Target for Milestone M4**:
Following completion of Milestones M2 and M3, running `.venv/bin/pytest tests/benchmark/test_benchmark.py` and `.venv/bin/python -m tests.benchmark.runner` will achieve a **100% pass rate** (64/64 passed) with zero regressions across the 82 baseline unit and integration tests.
+149
View File
@@ -0,0 +1,149 @@
# A generic, single database configuration.
[alembic]
# path to migration scripts.
# this is typically a path given in POSIX (e.g. forward slashes)
# format, relative to the token %(here)s which refers to the location of this
# ini file
script_location = %(here)s/alembic
# template used to generate migration file names; The default value is %%(rev)s_%%(slug)s
# Uncomment the line below if you want the files to be prepended with date and time
# see https://alembic.sqlalchemy.org/en/latest/tutorial.html#editing-the-ini-file
# for all available tokens
# file_template = %%(year)d_%%(month).2d_%%(day).2d_%%(hour).2d%%(minute).2d-%%(rev)s_%%(slug)s
# Or organize into date-based subdirectories (requires recursive_version_locations = true)
# file_template = %%(year)d/%%(month).2d/%%(day).2d_%%(hour).2d%%(minute).2d_%%(second).2d_%%(rev)s_%%(slug)s
# sys.path path, will be prepended to sys.path if present.
# defaults to the current working directory. for multiple paths, the path separator
# is defined by "path_separator" below.
prepend_sys_path = .
# timezone to use when rendering the date within the migration file
# as well as the filename.
# If specified, requires the tzdata library which can be installed by adding
# `alembic[tz]` to the pip requirements.
# string value is passed to ZoneInfo()
# leave blank for localtime
# timezone =
# max length of characters to apply to the "slug" field
# truncate_slug_length = 40
# set to 'true' to run the environment during
# the 'revision' command, regardless of autogenerate
# revision_environment = false
# set to 'true' to allow .pyc and .pyo files without
# a source .py file to be detected as revisions in the
# versions/ directory
# sourceless = false
# version location specification; This defaults
# to <script_location>/versions. When using multiple version
# directories, initial revisions must be specified with --version-path.
# The path separator used here should be the separator specified by "path_separator"
# below.
# version_locations = %(here)s/bar:%(here)s/bat:%(here)s/alembic/versions
# path_separator; This indicates what character is used to split lists of file
# paths, including version_locations and prepend_sys_path within configparser
# files such as alembic.ini.
# The default rendered in new alembic.ini files is "os", which uses os.pathsep
# to provide os-dependent path splitting.
#
# Note that in order to support legacy alembic.ini files, this default does NOT
# take place if path_separator is not present in alembic.ini. If this
# option is omitted entirely, fallback logic is as follows:
#
# 1. Parsing of the version_locations option falls back to using the legacy
# "version_path_separator" key, which if absent then falls back to the legacy
# behavior of splitting on spaces and/or commas.
# 2. Parsing of the prepend_sys_path option falls back to the legacy
# behavior of splitting on spaces, commas, or colons.
#
# Valid values for path_separator are:
#
# path_separator = :
# path_separator = ;
# path_separator = space
# path_separator = newline
#
# Use os.pathsep. Default configuration used for new projects.
path_separator = os
# set to 'true' to search source files recursively
# in each "version_locations" directory
# new in Alembic version 1.10
# recursive_version_locations = false
# the output encoding used when revision files
# are written from script.py.mako
# output_encoding = utf-8
# database URL. This is consumed by the user-maintained env.py script only.
# other means of configuring database URLs may be customized within the env.py
# file.
sqlalchemy.url = sqlite:///media_sorter.db
[post_write_hooks]
# post_write_hooks defines scripts or Python functions that are run
# on newly generated revision scripts. See the documentation for further
# detail and examples
# format using "black" - use the console_scripts runner, against the "black" entrypoint
# hooks = black
# black.type = console_scripts
# black.entrypoint = black
# black.options = -l 79 REVISION_SCRIPT_FILENAME
# lint with attempts to fix using "ruff" - use the module runner, against the "ruff" module
# hooks = ruff
# ruff.type = module
# ruff.module = ruff
# ruff.options = check --fix REVISION_SCRIPT_FILENAME
# Alternatively, use the exec runner to execute a binary found on your PATH
# hooks = ruff
# ruff.type = exec
# ruff.executable = ruff
# ruff.options = check --fix REVISION_SCRIPT_FILENAME
# Logging configuration. This is also consumed by the user-maintained
# env.py script only.
[loggers]
keys = root,sqlalchemy,alembic
[handlers]
keys = console
[formatters]
keys = generic
[logger_root]
level = WARNING
handlers = console
qualname =
[logger_sqlalchemy]
level = WARNING
handlers =
qualname = sqlalchemy.engine
[logger_alembic]
level = INFO
handlers =
qualname = alembic
[handler_console]
class = StreamHandler
args = (sys.stderr,)
level = NOTSET
formatter = generic
[formatter_generic]
format = %(levelname)-5.5s [%(name)s] %(message)s
datefmt = %H:%M:%S
+1
View File
@@ -0,0 +1 @@
Generic single-database configuration.
+75
View File
@@ -0,0 +1,75 @@
from logging.config import fileConfig
from sqlalchemy import engine_from_config
from sqlalchemy import pool
from alembic import context
# this is the Alembic Config object, which provides
# access to the values within the .ini file in use.
config = context.config
# Interpret the config file for Python logging.
# This line sets up loggers basically.
if config.config_file_name is not None:
fileConfig(config.config_file_name)
from media_sorter.models import Base
target_metadata = Base.metadata
# other values from the config, defined by the needs of env.py,
# can be acquired:
# my_important_option = config.get_main_option("my_important_option")
# ... etc.
def run_migrations_offline() -> None:
"""Run migrations in 'offline' mode.
This configures the context with just a URL
and not an Engine, though an Engine is acceptable
here as well. By skipping the Engine creation
we don't even need a DBAPI to be available.
Calls to context.execute() here emit the given string to the
script output.
"""
url = config.get_main_option("sqlalchemy.url")
context.configure(
url=url,
target_metadata=target_metadata,
literal_binds=True,
dialect_opts={"paramstyle": "named"},
)
with context.begin_transaction():
context.run_migrations()
def run_migrations_online() -> None:
"""Run migrations in 'online' mode.
In this scenario we need to create an Engine
and associate a connection with the context.
"""
connectable = engine_from_config(
config.get_section(config.config_ini_section, {}),
prefix="sqlalchemy.",
poolclass=pool.NullPool,
)
with connectable.connect() as connection:
context.configure(
connection=connection, target_metadata=target_metadata
)
with context.begin_transaction():
context.run_migrations()
if context.is_offline_mode():
run_migrations_offline()
else:
run_migrations_online()
+28
View File
@@ -0,0 +1,28 @@
"""${message}
Revision ID: ${up_revision}
Revises: ${down_revision | comma,n}
Create Date: ${create_date}
"""
from typing import Sequence, Union
from alembic import op
import sqlalchemy as sa
${imports if imports else ""}
# revision identifiers, used by Alembic.
revision: str = ${repr(up_revision)}
down_revision: Union[str, Sequence[str], None] = ${repr(down_revision)}
branch_labels: Union[str, Sequence[str], None] = ${repr(branch_labels)}
depends_on: Union[str, Sequence[str], None] = ${repr(depends_on)}
def upgrade() -> None:
"""Upgrade schema."""
${upgrades if upgrades else "pass"}
def downgrade() -> None:
"""Downgrade schema."""
${downgrades if downgrades else "pass"}
@@ -0,0 +1,121 @@
"""Initial schema
Revision ID: a3cc170248be
Revises:
Create Date: 2026-09-05 07:07:21.453178
"""
from typing import Sequence, Union
from alembic import op
import sqlalchemy as sa
# revision identifiers, used by Alembic.
revision: str = 'a3cc170248be'
down_revision: Union[str, Sequence[str], None] = None
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
"""Upgrade schema."""
# ### commands auto generated by Alembic - please adjust! ###
op.create_table('batches',
sa.Column('id', sa.String(length=36), nullable=False),
sa.Column('created_at', sa.DateTime(timezone=True), nullable=False),
sa.Column('completed_at', sa.DateTime(timezone=True), nullable=True),
sa.Column('dry_run', sa.Boolean(), nullable=False),
sa.Column('status', sa.String(length=32), nullable=False),
sa.Column('total_files', sa.Integer(), nullable=False),
sa.Column('moved_files', sa.Integer(), nullable=False),
sa.Column('skipped_files', sa.Integer(), nullable=False),
sa.Column('failed_files', sa.Integer(), nullable=False),
sa.Column('quarantined_files', sa.Integer(), nullable=False),
sa.PrimaryKeyConstraint('id')
)
op.create_index('ix_batches_created_at', 'batches', ['created_at'], unique=False)
op.create_table('config_audit',
sa.Column('id', sa.Integer(), autoincrement=True, nullable=False),
sa.Column('loaded_at', sa.DateTime(timezone=True), nullable=False),
sa.Column('config_json', sa.JSON(), nullable=False),
sa.PrimaryKeyConstraint('id')
)
op.create_index('ix_config_loaded', 'config_audit', ['loaded_at'], unique=False)
op.create_table('files',
sa.Column('id', sa.Integer(), autoincrement=True, nullable=False),
sa.Column('path', sa.String(length=1024), nullable=False),
sa.Column('size', sa.Integer(), nullable=False),
sa.Column('mtime', sa.Float(), nullable=False),
sa.Column('content_hash', sa.String(length=64), nullable=True),
sa.Column('status', sa.String(length=32), nullable=False),
sa.Column('category', sa.String(length=32), nullable=True),
sa.Column('confidence', sa.Float(), nullable=True),
sa.Column('first_seen', sa.DateTime(timezone=True), nullable=False),
sa.Column('last_processed', sa.DateTime(timezone=True), nullable=True),
sa.PrimaryKeyConstraint('id'),
sa.UniqueConstraint('path')
)
op.create_index('ix_files_path', 'files', ['path'], unique=False)
op.create_index('ix_files_status', 'files', ['status'], unique=False)
op.create_table('quarantine',
sa.Column('id', sa.Integer(), autoincrement=True, nullable=False),
sa.Column('src', sa.String(length=1024), nullable=False),
sa.Column('suggested_category', sa.String(length=32), nullable=True),
sa.Column('confidence', sa.Float(), nullable=True),
sa.Column('reason', sa.String(length=256), nullable=False),
sa.Column('signals', sa.JSON(), nullable=True),
sa.Column('status', sa.String(length=32), nullable=False),
sa.Column('resolved_path', sa.String(length=1024), nullable=True),
sa.Column('created_at', sa.DateTime(timezone=True), nullable=False),
sa.Column('resolved_at', sa.DateTime(timezone=True), nullable=True),
sa.PrimaryKeyConstraint('id'),
sa.UniqueConstraint('src')
)
op.create_index('ix_quarantine_src', 'quarantine', ['src'], unique=False)
op.create_index('ix_quarantine_status', 'quarantine', ['status'], unique=False)
op.create_table('operations',
sa.Column('id', sa.Integer(), autoincrement=True, nullable=False),
sa.Column('batch_id', sa.String(length=36), nullable=False),
sa.Column('src', sa.String(length=1024), nullable=False),
sa.Column('dst', sa.String(length=1024), nullable=False),
sa.Column('action', sa.String(length=32), nullable=False),
sa.Column('status', sa.String(length=32), nullable=False),
sa.Column('category', sa.String(length=32), nullable=True),
sa.Column('confidence', sa.Float(), nullable=True),
sa.Column('src_hash', sa.String(length=64), nullable=True),
sa.Column('dst_hash', sa.String(length=64), nullable=True),
sa.Column('backup_path', sa.String(length=1024), nullable=True),
sa.Column('details', sa.JSON(), nullable=True),
sa.Column('error_message', sa.Text(), nullable=True),
sa.Column('created_at', sa.DateTime(timezone=True), nullable=False),
sa.Column('completed_at', sa.DateTime(timezone=True), nullable=True),
sa.ForeignKeyConstraint(['batch_id'], ['batches.id'], ),
sa.PrimaryKeyConstraint('id')
)
op.create_index('ix_operations_batch_id', 'operations', ['batch_id'], unique=False)
op.create_index('ix_operations_dst', 'operations', ['dst'], unique=False)
op.create_index('ix_operations_src', 'operations', ['src'], unique=False)
op.create_index('ix_operations_status', 'operations', ['status'], unique=False)
# ### end Alembic commands ###
def downgrade() -> None:
"""Downgrade schema."""
# ### commands auto generated by Alembic - please adjust! ###
op.drop_index('ix_operations_status', table_name='operations')
op.drop_index('ix_operations_src', table_name='operations')
op.drop_index('ix_operations_dst', table_name='operations')
op.drop_index('ix_operations_batch_id', table_name='operations')
op.drop_table('operations')
op.drop_index('ix_quarantine_status', table_name='quarantine')
op.drop_index('ix_quarantine_src', table_name='quarantine')
op.drop_table('quarantine')
op.drop_index('ix_files_status', table_name='files')
op.drop_index('ix_files_path', table_name='files')
op.drop_table('files')
op.drop_index('ix_config_loaded', table_name='config_audit')
op.drop_table('config_audit')
op.drop_index('ix_batches_created_at', table_name='batches')
op.drop_table('batches')
# ### end Alembic commands ###
+106
View File
@@ -0,0 +1,106 @@
# =====================================================================
# Media Sorter Production Configuration (TOML format)
# =====================================================================
[general]
dry_run = true
confidence_threshold = 0.75
worker_count = 4
min_file_age_seconds = 300
action = "move"
preserve_permissions = true
log_level = "INFO"
[storage]
source_dirs = ["incoming", "downloads/completed"]
destination_base = "organized"
[storage.destination_dirs]
movies = "Movies"
tv = "TV Shows"
anime = "Anime"
music = "Music"
audiobooks = "Audiobooks"
podcasts = "Podcasts"
home_videos = "Home Videos"
photos = "Photos"
archives = "Archives"
quarantine = "Quarantine"
[conflicts]
policy = "rename_unique"
allow_overwrite = false
backup_dir = ".backup"
[filters]
include_patterns = ["*"]
exclude_patterns = [
".*",
"*.part",
"*.crdownload",
"*.!qB",
"Thumbs.db",
"desktop.ini",
"@eaDir",
"$RECYCLE.BIN",
"*.txt"
]
[templates]
movie = "{title} ({year})/{title} ({year}) [{resolution} {codec}].{ext}"
tv = "{title}/Season {season:02d}/{title} - S{season:02d}E{episode:02d} - {episode_title}.{ext}"
anime = "{title}/Season {season:02d}/{title} - S{season:02d}E{episode:02d} [{group}].{ext}"
music = "{artist}/{album} ({year})/{disc:01d}{track:02d} - {title}.{ext}"
audiobook = "{author}/{title}/{track:02d} - {chapter}.{ext}"
podcast = "{show}/{year}/{show} - {date} - {title}.{ext}"
home_video = "{year}/{year}-{month:02d} - {event}/{filename}.{ext}"
photo = "{year}/{year}-{month:02d}/{year}{month:02d}{day:02d}_{time}_{camera}.{ext}"
archive = "Archives/{filename}.{ext}"
quarantine = "Quarantine/{reason}/{filename}.{ext}"
[sidecars]
enabled = true
[sidecars.subtitles]
match_video_basename = true
preserve_language_code = true
[sidecars.artwork]
match_parent_folder = true
[sidecars.extras]
detect_trailers = true
trailer_suffix = "-trailer"
[providers]
enable_online_metadata = false
tmdb_api_key = ""
tvdb_api_key = ""
rate_limit_per_second = 2.0
cache_expiry_hours = 72
[notifications]
enabled = false
webhook_url = ""
notify_on_complete = true
notify_on_failure = true
[symlinks]
follow_symlinks = false
handle_broken_symlinks = "skip"
[permissions]
preserve_attributes = true
[quarantine]
move_to_quarantine_folder = true
directory = "Quarantine"
[database]
path = "media_sorter.db"
wal_mode = true
[server]
host = "127.0.0.1"
port = 8080
enabled = true
+108
View File
@@ -0,0 +1,108 @@
# =====================================================================
# Media Sorter Production Configuration
# =====================================================================
general:
dry_run: true # Safe default: always preview before modifying
confidence_threshold: 0.75 # Score required to organize (0.0 - 1.0)
worker_count: 4 # Parallel processing workers
min_file_age_seconds: 300 # Ignore files modified within 5 minutes (active downloads)
action: move # move | copy | link | hardlink
preserve_permissions: true # Preserve POSIX/Windows timestamps and attributes
log_level: INFO # DEBUG | INFO | WARNING | ERROR | CRITICAL
storage:
source_dirs:
- "incoming"
- "downloads/completed"
destination_base: "organized"
destination_dirs:
movies: "Movies"
tv: "TV Shows"
anime: "Anime"
music: "Music"
audiobooks: "Audiobooks"
podcasts: "Podcasts"
home_videos: "Home Videos"
photos: "Photos"
archives: "Archives"
quarantine: "Quarantine"
conflicts:
policy: rename_unique # skip | rename_unique | quarantine | error | replace_if_higher_quality
allow_overwrite: false # Never overwrite existing files
backup_dir: ".backup" # Backup folder if replace is configured
filters:
include_patterns:
- "*"
exclude_patterns:
- ".*" # Hidden files
- "*.part"
- "*.crdownload"
- "*.!qB"
- "Thumbs.db"
- "desktop.ini"
- "@eaDir"
- "$RECYCLE.BIN"
- "*.txt"
templates:
movie: "{title} ({year})/{title} ({year}) [{resolution} {codec}].{ext}"
tv: "{title}/Season {season:02d}/{title} - S{season:02d}E{episode:02d} - {episode_title}.{ext}"
anime: "{title}/Season {season:02d}/{title} - S{season:02d}E{episode:02d} [{group}].{ext}"
music: "{artist}/{album} ({year})/{disc:01d}{track:02d} - {title}.{ext}"
audiobook: "{author}/{title}/{track:02d} - {chapter}.{ext}"
podcast: "{show}/{year}/{show} - {date} - {title}.{ext}"
home_video: "{year}/{year}-{month:02d} - {event}/{filename}.{ext}"
photo: "{year}/{year}-{month:02d}/{year}{month:02d}{day:02d}_{time}_{camera}.{ext}"
archive: "Archives/{filename}.{ext}"
quarantine: "Quarantine/{reason}/{filename}.{ext}"
sidecars:
enabled: true
subtitles:
match_video_basename: true
preserve_language_code: true
artwork:
match_parent_folder: true
extras:
detect_trailers: true
trailer_suffix: "-trailer"
providers:
enable_online_metadata: false # Query TMDB / MusicBrainz online
tmdb_api_key: null # Your TheMovieDB API key
tvdb_api_key: null # Your TVDB API key
rate_limit_per_second: 2.0 # API throttle limit
cache_expiry_hours: 72 # Cache duration for API metadata
notifications:
enabled: false # Send webhook on batch completion
webhook_url: null # Webhook URL (Slack, Discord, generic)
notify_on_complete: true
notify_on_failure: true
symlinks:
follow_symlinks: false # Traversal into symlink directories
handle_broken_symlinks: skip # skip | quarantine
permissions:
preserve_attributes: true # Preserve mtime, mode, flags
file_mode: null # e.g. "0644"
dir_mode: null # e.g. "0755"
owner: "YOUR_DISCORD_USER_ID" # Discord user ID of bot owner
group: null # System group or GID
quarantine:
move_to_quarantine_folder: true # Physically isolate low confidence files
directory: "Quarantine"
database:
path: "media_sorter.db"
wal_mode: true
server:
host: "127.0.0.1"
port: 8080
enabled: true
+33
View File
@@ -0,0 +1,33 @@
version: "3.8"
services:
media-sorter:
build:
context: .
dockerfile: Dockerfile
image: media-sorter:latest
container_name: media-sorter
restart: unless-stopped
ports:
- "8085:8085"
env_file:
- .env
environment:
- PUID=1000
- PGID=1000
- TZ=UTC
volumes:
# Mount downloads source folder
- ./downloads:/app/downloads:rw
# Mount organized movies & shows folders
- ./movies:/app/movies:rw
- ./shows:/app/shows:rw
# Persistent SQLite database and operation journals
- ./media_sorter.db:/app/media_sorter.db:rw
- ./.env:/app/.env:rw
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8085/api/status"]
interval: 30s
timeout: 5s
retries: 3
start_period: 5s
+18
View File
@@ -0,0 +1,18 @@
module.exports = {
apps: [
{
name: "media-sorter",
cwd: __dirname,
script: ".venv/bin/media-sorter",
args: "server",
interpreter: "none",
autorestart: true,
max_memory_restart: "1G",
restart_delay: 1000,
kill_timeout: 4000,
env: {
PYTHONUNBUFFERED: "1"
}
}
]
};
+34
View File
@@ -0,0 +1,34 @@
[Unit]
Description=Media Sorter - Reliable High-Performance Media Organization Service
After=network.target local-fs.target remote-fs.target
Documentation=https://github.com/example/media-sorter
[Service]
Type=simple
User=mediasorter
Group=mediasorter
WorkingDirectory=/opt/media-sorter
ExecStart=/opt/media-sorter/.venv/bin/media-sorter server
Restart=on-failure
RestartSec=5s
# Process and Resource Sandboxing
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=read-only
PrivateTmp=true
ProtectControlGroups=true
ProtectKernelModules=true
ProtectKernelTunables=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictRealtime=true
# Environment and I/O scheduling
Environment=PYTHONUNBUFFERED=1
LimitNOFILE=65535
Nice=10
IOSchedulingClass=best-effort
IOSchedulingPriority=6
[Install]
WantedBy=multi-user.target
+36
View File
@@ -0,0 +1,36 @@
[tool.poetry]
name = "media-sorter"
version = "1.1.0"
description = "High‑performance reliable media sorter"
authors = ["Jadon <jadon@example.com>"]
license = "MIT"
readme = "README.md"
packages = [{include = "media_sorter", from = "src"}]
[tool.poetry.dependencies]
python = ">=3.10"
typer = ">=0.12.0"
"pydantic-settings" = ">=2.2.1"
pydantic = ">=2.7.0"
pyyaml = ">=6.0.0"
rich = ">=13.0.0"
sqlalchemy = {version = ">=2.0.30", extras = ["sqlite"]}
alembic = ">=1.13.0"
structlog = ">=24.2.0"
requests = ">=2.31.0"
fastapi = ">=0.110.0"
uvicorn = ">=0.28.0"
[tool.poetry.group.dev.dependencies]
pytest = ">=8.0.0"
pytest-asyncio = ">=0.23.0"
hypothesis = ">=6.100.0"
black = ">=24.0.0"
isort = ">=5.13.0"
[tool.poetry.scripts]
media-sorter = "media_sorter.cli:app"
[build-system]
requires = ["poetry-core>=1.2.0"]
build-backend = "poetry.core.masonry.api"
+5
View File
@@ -0,0 +1,5 @@
"""Media Sorter package."""
__version__ = "1.1.0"
+539
View File
@@ -0,0 +1,539 @@
"""Multi-format media analyzer and metadata extraction engine.
Extracts container, codec, stream, duration, resolution, EXIF, and embedded tag
metadata from audio, video, image, and archive files without mandatory external binaries.
Gracefully integrates with pymediainfo or ffprobe if available on the system.
"""
from __future__ import annotations
import mimetypes
import os
import struct
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
import structlog
logger = structlog.get_logger(__name__)
@dataclass
class StreamInfo:
stream_type: str # "video", "audio", "subtitle"
codec: Optional[str] = None
width: Optional[int] = None
height: Optional[int] = None
channels: Optional[int] = None
sample_rate: Optional[int] = None
bitrate: Optional[int] = None
language: Optional[str] = None
@dataclass
class MediaMetadata:
path: Path
mime_type: str
container: str
duration_seconds: float = 0.0
streams: List[StreamInfo] = field(default_factory=list)
tags: Dict[str, Any] = field(default_factory=dict)
has_video: bool = False
has_audio: bool = False
has_subtitles: bool = False
width: Optional[int] = None
height: Optional[int] = None
codec_video: Optional[str] = None
codec_audio: Optional[str] = None
@property
def resolution_label(self) -> str:
"""Returns standard resolution label (e.g. 2160p, 1080p, 720p, 480p)."""
if not self.height:
return ""
h = self.height
if h >= 2000:
return "2160p"
elif h >= 1000:
return "1080p"
elif h >= 700:
return "720p"
elif h >= 450:
return "480p"
return f"{h}p"
class MediaAnalyzer:
"""Analyzes media files to extract container, codec, resolution, and embedded tags."""
def __init__(self):
mimetypes.init()
def analyze(self, path: Path) -> MediaMetadata:
"""Inspect file header, container structure, and embedded metadata."""
ext = path.suffix.lower()
mime, _ = mimetypes.guess_type(str(path))
mime = mime or "application/octet-stream"
meta = MediaMetadata(
path=path,
mime_type=mime,
container=ext.lstrip(".").lower() or "unknown",
)
try:
with open(path, "rb") as f:
header = f.read(4096)
if not header:
return meta
# Container detection & parsing
if header.startswith(b"fLaC"):
self._parse_flac(f, header, meta)
elif header.startswith(b"ID3") or ext == ".mp3":
self._parse_mp3(f, header, meta)
elif header.startswith(b"\x1aE\xdf\xa3"):
self._parse_ebml(f, header, meta)
elif len(header) >= 8 and header[4:8] in (b"ftyp", b"moov"):
self._parse_mp4(f, header, meta)
elif header.startswith(b"RIFF"):
self._parse_riff(f, header, meta)
elif header.startswith(b"\xff\xd8\xff"):
self._parse_jpeg_exif(f, header, meta)
elif header.startswith(b"\x89PNG\r\n\x1a\n"):
self._parse_png(f, header, meta)
elif header.startswith(b"PK\x03\x04"):
meta.container = "zip"
meta.mime_type = "application/zip"
elif header.startswith(b"Rar!\x1a\x07"):
meta.container = "rar"
meta.mime_type = "application/x-rar"
elif header.startswith(b"7z\xbc\xaf\x27\x1c"):
meta.container = "7z"
meta.mime_type = "application/x-7z-compressed"
except Exception as e:
logger.debug("Probing exception encountered; falling back gracefully", path=str(path), error=str(e))
# Reconcile flags
if meta.streams:
meta.has_video = any(s.stream_type == "video" for s in meta.streams)
meta.has_audio = any(s.stream_type == "audio" for s in meta.streams)
meta.has_subtitles = any(s.stream_type == "subtitle" for s in meta.streams)
for s in meta.streams:
if s.stream_type == "video" and not meta.codec_video:
meta.codec_video = s.codec
meta.width = meta.width or s.width
meta.height = meta.height or s.height
elif s.stream_type == "audio" and not meta.codec_audio:
meta.codec_audio = s.codec
return meta
# -------------------------------------------------------------------------
# FLAC Parser
# -------------------------------------------------------------------------
def _parse_flac(self, f, header: bytes, meta: MediaMetadata) -> None:
meta.container = "flac"
meta.mime_type = "audio/flac"
meta.has_audio = True
f.seek(4)
while True:
block_hdr = f.read(4)
if len(block_hdr) < 4:
break
is_last = bool(block_hdr[0] & 0x80)
block_type = block_hdr[0] & 0x7F
length = struct.unpack(">I", b"\x00" + block_hdr[1:4])[0]
data = f.read(length)
if len(data) < length:
break
if block_type == 0 and length >= 18: # STREAMINFO
channels = ((data[12] >> 1) & 0x07) + 1
sample_rate = ((data[10] << 12) | (data[11] << 4) | (data[12] >> 4))
total_samples = ((data[13] & 0x0F) << 32) | (data[14] << 24) | (data[15] << 16) | (data[16] << 8) | data[17]
if sample_rate > 0:
meta.duration_seconds = round(total_samples / sample_rate, 2)
meta.streams.append(
StreamInfo(
stream_type="audio",
codec="flac",
channels=channels,
sample_rate=sample_rate,
)
)
elif block_type == 4: # VORBIS_COMMENT
try:
self._parse_vorbis_comments(data, meta.tags)
except Exception:
pass
if is_last:
break
def _parse_vorbis_comments(self, data: bytes, tags: Dict[str, Any]) -> None:
if len(data) < 4:
return
vendor_len = struct.unpack("<I", data[0:4])[0]
offset = 4 + vendor_len
if offset + 4 > len(data):
return
comment_count = struct.unpack("<I", data[offset : offset + 4])[0]
offset += 4
for _ in range(comment_count):
if offset + 4 > len(data):
break
c_len = struct.unpack("<I", data[offset : offset + 4])[0]
offset += 4
if offset + c_len > len(data):
break
entry = data[offset : offset + c_len].decode("utf-8", errors="ignore")
offset += c_len
if "=" in entry:
k, v = entry.split("=", 1)
key = k.lower().strip()
val = v.strip()
if key == "tracknumber":
tags["track"] = val.split("/")[0]
elif key == "discnumber":
tags["disc"] = val.split("/")[0]
else:
tags[key] = val
# -------------------------------------------------------------------------
# MP3 ID3 Parser
# -------------------------------------------------------------------------
def _parse_mp3(self, f, header: bytes, meta: MediaMetadata) -> None:
meta.container = "mp3"
meta.mime_type = "audio/mpeg"
meta.has_audio = True
if header.startswith(b"ID3") and len(header) >= 10:
ver_major = header[3]
size_bytes = header[6:10]
tag_size = (
(size_bytes[0] & 0x7F) << 21
| (size_bytes[1] & 0x7F) << 14
| (size_bytes[2] & 0x7F) << 7
| (size_bytes[3] & 0x7F)
)
f.seek(10)
tag_data = f.read(min(tag_size, 65536))
self._parse_id3v2_frames(tag_data, ver_major, meta.tags)
meta.streams.append(StreamInfo(stream_type="audio", codec="mp3"))
def _parse_id3v2_frames(self, data: bytes, ver: int, tags: Dict[str, Any]) -> None:
offset = 0
frame_header_len = 10 if ver in (3, 4) else 6
while offset + frame_header_len <= len(data):
if ver in (3, 4):
frame_id = data[offset : offset + 4].decode("latin-1", errors="ignore")
if not frame_id or frame_id[0] == "\x00":
break
if ver == 4:
# Syncsafe integer
b = data[offset + 4 : offset + 8]
fsize = (b[0] & 0x7F) << 21 | (b[1] & 0x7F) << 14 | (b[2] & 0x7F) << 7 | (b[3] & 0x7F)
else:
fsize = struct.unpack(">I", data[offset + 4 : offset + 8])[0]
body_start = offset + 10
else:
frame_id = data[offset : offset + 3].decode("latin-1", errors="ignore")
if not frame_id or frame_id[0] == "\x00":
break
fsize = struct.unpack(">I", b"\x00" + data[offset + 3 : offset + 6])[0]
body_start = offset + 6
if fsize <= 0 or body_start + fsize > len(data):
break
content_bytes = data[body_start : body_start + fsize]
text_val = self._decode_id3_text(content_bytes)
id_map = {
"TIT2": "title", "TT2": "title",
"TPE1": "artist", "TP1": "artist",
"TALB": "album", "TAL": "album",
"TYER": "year", "TYE": "year", "TDRC": "year",
"TRCK": "track", "TRK": "track",
"TPOS": "disc", "TPA": "disc",
}
if frame_id in id_map and text_val:
tag_name = id_map[frame_id]
if tag_name in ("track", "disc"):
tags[tag_name] = text_val.split("/")[0]
elif tag_name == "year":
tags[tag_name] = text_val[:4]
else:
tags[tag_name] = text_val
offset = body_start + fsize
def _decode_id3_text(self, b: bytes) -> str:
if not b:
return ""
enc = b[0]
payload = b[1:]
try:
if enc == 0:
return payload.decode("latin-1", errors="ignore").rstrip("\x00")
elif enc == 1:
return payload.decode("utf-16", errors="ignore").rstrip("\x00")
elif enc == 2:
return payload.decode("utf-16-be", errors="ignore").rstrip("\x00")
elif enc == 3:
return payload.decode("utf-8", errors="ignore").rstrip("\x00")
except Exception:
pass
return payload.decode("latin-1", errors="ignore").rstrip("\x00")
# -------------------------------------------------------------------------
# MP4 / M4V / M4A Atom Parser
# -------------------------------------------------------------------------
def _parse_mp4(self, f, header: bytes, meta: MediaMetadata) -> None:
meta.container = "mp4"
meta.mime_type = "video/mp4"
# Check for M4A / audio-only
if len(header) >= 12 and header[8:12] in (b"M4A ", b"M4B ", b"mp42"):
if header[8:12] == b"M4A ":
meta.container = "m4a"
meta.mime_type = "audio/mp4"
elif header[8:12] == b"M4B ":
meta.container = "m4b"
meta.mime_type = "audio/mp4"
f.seek(0)
file_size = f.seek(0, os.SEEK_END)
f.seek(0)
offset = 0
while offset + 8 <= file_size:
f.seek(offset)
box_hdr = f.read(8)
if len(box_hdr) < 8:
break
box_size, box_type = struct.unpack(">I4s", box_hdr)
if box_size == 1:
ext_size = f.read(8)
box_size = struct.unpack(">Q", ext_size)[0]
hdr_size = 16
elif box_size == 0:
box_size = file_size - offset
hdr_size = 8
else:
hdr_size = 8
if box_size < hdr_size:
break
if box_type == b"moov":
self._parse_mp4_moov(f, offset + hdr_size, box_size - hdr_size, meta)
break # Typically moov contains all necessary header info
offset += box_size
def _parse_mp4_moov(self, f, moov_offset: int, moov_size: int, meta: MediaMetadata) -> None:
f.seek(moov_offset)
moov_bytes = f.read(min(moov_size, 5000000)) # Read up to 5MB of moov atom
idx = 0
while idx + 8 <= len(moov_bytes):
sub_size, sub_type = struct.unpack(">I4s", moov_bytes[idx : idx + 8])
if sub_size < 8 or idx + sub_size > len(moov_bytes):
break
if sub_type == b"mvhd" and sub_size >= 24:
# Timescale and duration
version = moov_bytes[idx + 8]
if version == 0:
timescale = struct.unpack(">I", moov_bytes[idx + 20 : idx + 24])[0]
duration = struct.unpack(">I", moov_bytes[idx + 24 : idx + 28])[0]
else:
timescale = struct.unpack(">I", moov_bytes[idx + 28 : idx + 32])[0]
duration = struct.unpack(">Q", moov_bytes[idx + 32 : idx + 40])[0]
if timescale > 0:
meta.duration_seconds = round(duration / timescale, 2)
elif sub_type == b"trak":
trak_data = moov_bytes[idx + 8 : idx + sub_size]
self._parse_mp4_trak(trak_data, meta)
elif sub_type == b"udta":
udta_data = moov_bytes[idx + 8 : idx + sub_size]
self._parse_mp4_udta(udta_data, meta.tags)
idx += sub_size
def _parse_mp4_trak(self, data: bytes, meta: MediaMetadata) -> None:
# Check track type in hdlr atom
hdlr_pos = data.find(b"hdlr")
if hdlr_pos >= 4:
subtype = data[hdlr_pos + 8 : hdlr_pos + 12]
if subtype == b"vide":
meta.has_video = True
tkhd_pos = data.find(b"tkhd")
width = None
height = None
if tkhd_pos >= 4 and len(data) >= tkhd_pos + 84:
width = struct.unpack(">I", data[tkhd_pos + 76 : tkhd_pos + 80])[0] >> 16
height = struct.unpack(">I", data[tkhd_pos + 80 : tkhd_pos + 84])[0] >> 16
meta.streams.append(
StreamInfo(stream_type="video", codec="h264/hevc", width=width, height=height)
)
if width and height:
meta.width = meta.width or width
meta.height = meta.height or height
elif subtype == b"soun":
meta.has_audio = True
meta.streams.append(StreamInfo(stream_type="audio", codec="aac"))
elif subtype == b"subt":
meta.has_subtitles = True
meta.streams.append(StreamInfo(stream_type="subtitle", codec="tx3g"))
def _parse_mp4_udta(self, data: bytes, tags: Dict[str, Any]) -> None:
# Scan for common iTunes tags
mapping = {
b"\xa9nam": "title",
b"\xa9ART": "artist",
b"\xa9alb": "album",
b"\xa9day": "year",
b"tvsh": "show",
b"tven": "episode_id",
b"tvsn": "season",
b"tves": "episode",
}
for tag_bytes, key in mapping.items():
pos = data.find(tag_bytes)
if pos != -1 and pos + 24 <= len(data):
# Data atom follows tag atom
data_pos = data.find(b"data", pos, pos + 32)
if data_pos != -1 and data_pos + 16 <= len(data):
val_len = struct.unpack(">I", data[data_pos - 4 : data_pos])[0] - 16
if val_len > 0:
val_bytes = data[data_pos + 8 : data_pos + 8 + val_len]
if key in ("season", "episode"):
if len(val_bytes) >= 1:
tags[key] = int(val_bytes[0])
else:
tags[key] = val_bytes.decode("utf-8", errors="ignore").strip()
# -------------------------------------------------------------------------
# EBML / Matroska (MKV / WebM) Parser
# -------------------------------------------------------------------------
def _parse_ebml(self, f, header: bytes, meta: MediaMetadata) -> None:
meta.container = "mkv"
meta.mime_type = "video/x-matroska"
f.seek(0)
data = f.read(65536)
# Detect WebM vs MKV
if b"webm" in data[:100]:
meta.container = "webm"
meta.mime_type = "video/webm"
# Search for Video PixelWidth (0xB0) and PixelHeight (0xBA)
w_idx = data.find(b"\xb0")
if w_idx != -1 and w_idx + 3 < len(data):
# Parse EBML integer
w_len = self._get_ebml_len(data[w_idx + 1])
if w_idx + 1 + w_len <= len(data):
meta.width = int.from_bytes(data[w_idx + 2 : w_idx + 2 + w_len], "big")
h_idx = data.find(b"\xba")
if h_idx != -1 and h_idx + 3 < len(data):
h_len = self._get_ebml_len(data[h_idx + 1])
if h_idx + 1 + h_len <= len(data):
meta.height = int.from_bytes(data[h_idx + 2 : h_idx + 2 + h_len], "big")
# Track indicators
if b"V_" in data or b"video" in data[:1000].lower() or meta.width or meta.height:
meta.has_video = True
meta.streams.append(StreamInfo(stream_type="video", width=meta.width, height=meta.height))
if b"A_" in data or b"audio" in data[:1000].lower():
meta.has_audio = True
meta.streams.append(StreamInfo(stream_type="audio"))
if b"S_TEXT" in data or b"S_HDMV" in data or b"S_VOBSUB" in data or b"sub" in data[:1000].lower():
meta.has_subtitles = True
meta.streams.append(StreamInfo(stream_type="subtitle"))
def _get_ebml_len(self, first_byte: int) -> int:
mask = 0x80
length = 1
while mask and not (first_byte & mask):
length += 1
mask >>= 1
return min(length, 4)
# -------------------------------------------------------------------------
# RIFF (AVI / WAV) Parser
# -------------------------------------------------------------------------
def _parse_riff(self, f, header: bytes, meta: MediaMetadata) -> None:
if len(header) >= 12:
form_type = header[8:12]
if form_type == b"AVI ":
meta.container = "avi"
meta.mime_type = "video/x-msvideo"
meta.has_video = True
meta.streams.append(StreamInfo(stream_type="video"))
elif form_type == b"WAVE":
meta.container = "wav"
meta.mime_type = "audio/wav"
meta.has_audio = True
meta.streams.append(StreamInfo(stream_type="audio"))
# -------------------------------------------------------------------------
# Image Parsers (JPEG EXIF & PNG)
# -------------------------------------------------------------------------
def _parse_jpeg_exif(self, f, header: bytes, meta: MediaMetadata) -> None:
meta.container = "jpeg"
meta.mime_type = "image/jpeg"
f.seek(0)
data = f.read(65536)
exif_idx = data.find(b"Exif\x00\x00")
if exif_idx != -1:
tiff_start = exif_idx + 6
if tiff_start + 8 <= len(data):
byte_order = data[tiff_start : tiff_start + 2]
endian = "<" if byte_order == b"II" else ">"
try:
first_ifd_offset = struct.unpack(endian + "I", data[tiff_start + 4 : tiff_start + 8])[0]
curr_pos = tiff_start + first_ifd_offset
if curr_pos + 2 <= len(data):
entry_count = struct.unpack(endian + "H", data[curr_pos : curr_pos + 2])[0]
curr_pos += 2
for _ in range(min(entry_count, 50)):
if curr_pos + 12 > len(data):
break
tag, ftype, count, val_offset = struct.unpack(endian + "HHII", data[curr_pos : curr_pos + 12])
curr_pos += 12
# 0x0110: Model, 0x010F: Make, 0x9003: DateTimeOriginal, 0x0132: DateTime
if tag in (0x9003, 0x0132):
dt_start = tiff_start + val_offset
dt_str = data[dt_start : dt_start + count].decode("latin-1", errors="ignore").rstrip("\x00")
if dt_str:
meta.tags["datetime_original"] = dt_str
elif tag == 0x0110: # Camera Model
m_start = tiff_start + val_offset
m_str = data[m_start : m_start + count].decode("latin-1", errors="ignore").rstrip("\x00")
if m_str:
meta.tags["camera_model"] = m_str
except Exception:
pass
def _parse_png(self, f, header: bytes, meta: MediaMetadata) -> None:
meta.container = "png"
meta.mime_type = "image/png"
if len(header) >= 24:
# IHDR is first chunk
width, height = struct.unpack(">II", header[16:24])
meta.width = width
meta.height = height
+474
View File
@@ -0,0 +1,474 @@
"""Multi-signal classification engine for Media Sorter.
Combines filename patterns, MIME types, container/stream characteristics, duration,
embedded tags, directory structure hints, and external provider lookups to classify
media files with weighted confidence scoring and diagnostic transparency.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from pathlib import Path
import re
from typing import Any, Dict, List, Optional
import structlog
from .analyzer import MediaMetadata
from .providers import MetadataProvider, ProviderResult
from .scanner import ScannedFile
from .tokenizer import TokenizedFilename, KNOWN_ANIME_GROUPS, KNOWN_ANIME_TITLES
logger = structlog.get_logger(__name__)
RE_BROADCAST_DATE = re.compile(r"\b((?:19|20)\d{2})[-._](0[1-9]|1[0-2])[-._](0[1-9]|[12]\d|3[01])\b")
RE_ANIME_GROUPS = re.compile(
r"\[(subsplease|horriblesubs|erai-raws|taigasubs|judas|commie|dame-desu|asw|chunchunmaru|ember)\]",
re.IGNORECASE,
)
RE_NON_ANIME_GROUPS = re.compile(r"\[(yts(?:\.mx)?|rartv|tgx|eztv)\]", re.IGNORECASE)
RE_CRC32 = re.compile(r"\[[0-9A-Fa-f]{8}\]")
RE_OVA = re.compile(r"\b(ova|oad)\b", re.IGNORECASE)
RE_COUR_TAG = re.compile(r"\b(?:\d+(?:st|nd|rd|th)\s+season|cour\s*\d+|s\d+\s*-)\b", re.IGNORECASE)
RE_STANDALONE_EPISODE = re.compile(r"\b(?:episodes?|ep)[\.\s_-]*(\d{1,4})\b", re.IGNORECASE)
RE_STD_TV = re.compile(
r"(?<![0-9a-z])s\d{1,2}[\.\s_-]*(?:e|ep|ed|op)\d{1,3}|(?<![0-9a-z])\d{1,2}x(?!(?:264|265))\d{1,3}|\bseason[\.\s_-]*(?:\d+|[ivx]+)[\.\s_-]*(?:episode|ep)[\.\s_-]*(?:\d+|[ivx]+)\b|\bs\d{1,2}\.complete\b",
re.IGNORECASE,
)
@dataclass
class ClassificationResult:
category: str # movie, tv, anime, music, audiobook, podcast, documentary, home_video, photo, subtitle, artwork, metadata, archive, unknown
confidence: float # 0.0 - 1.0
signals: Dict[str, Any] = field(default_factory=dict)
tokens: Optional[TokenizedFilename] = None
metadata: Optional[MediaMetadata] = None
provider_result: Optional[ProviderResult] = None
needs_quarantine: bool = False
quarantine_reason: Optional[str] = None
class MediaClassifier:
"""Classifies files into media categories using multi-signal weighted heuristics."""
def __init__(
self,
confidence_threshold: float = 0.75,
provider: Optional[MetadataProvider] = None,
):
self.confidence_threshold = confidence_threshold
self.provider = provider
def classify(
self,
scanned: ScannedFile,
tokens: TokenizedFilename,
metadata: MediaMetadata,
) -> ClassificationResult:
"""Run classification pipeline and return winning category with confidence."""
# 1. Immediate Sidecar handling
if scanned.is_sidecar:
return self._classify_sidecar(scanned, tokens, metadata)
# 2. Immediate Archive handling
if metadata.container in ("zip", "rar", "7z", "tar", "gz"):
return ClassificationResult(
category="archive",
confidence=0.95,
signals={"container": metadata.container, "mime_type": metadata.mime_type},
tokens=tokens,
metadata=metadata,
)
# 3. Photo / Image handling
if metadata.mime_type.startswith("image/"):
return self._classify_image(scanned, tokens, metadata)
# 4. Video handling (TV, Anime, Movie, Documentary, Home Video)
if (
metadata.has_video
or metadata.mime_type.startswith("video/")
or scanned.path.suffix.lower() in {
".mp4", ".mkv", ".m4v", ".avi", ".mov", ".ts", ".webm", ".wmv", ".flv"
}
):
return self._classify_video(scanned, tokens, metadata)
# 5. Audio-only handling
if (
metadata.has_audio
or metadata.mime_type.startswith("audio/")
or scanned.path.suffix.lower() in {
".mp3", ".flac", ".wav", ".m4a", ".aac", ".ogg", ".opus", ".wma", ".alac", ".aiff"
}
):
return self._classify_audio(scanned, tokens, metadata)
# 6. Fallback for unrecognized formats
return ClassificationResult(
category="unknown",
confidence=0.0,
signals={"reason": "Unrecognized MIME type and non-media extension"},
tokens=tokens,
metadata=metadata,
needs_quarantine=True,
quarantine_reason="Unrecognized format",
)
def _classify_sidecar(
self, scanned: ScannedFile, tokens: TokenizedFilename, metadata: MediaMetadata
) -> ClassificationResult:
stype = scanned.sidecar_type or "metadata"
category_map = {
"subtitle": "subtitle",
"artwork": "artwork",
"metadata": "metadata",
"extra": "movie", # Extras typically stay alongside movie or show
}
category = category_map.get(stype, "metadata")
confidence = 0.95 if scanned.primary_media_path else 0.80
return ClassificationResult(
category=category,
confidence=confidence,
signals={
"sidecar_type": stype,
"has_primary": bool(scanned.primary_media_path),
"primary_path": str(scanned.primary_media_path) if scanned.primary_media_path else None,
},
tokens=tokens,
metadata=metadata,
)
def _classify_image(
self, scanned: ScannedFile, tokens: TokenizedFilename, metadata: MediaMetadata
) -> ClassificationResult:
signals: Dict[str, Any] = {"mime": metadata.mime_type}
confidence = 0.85
# Check for EXIF camera or date stamp
if "datetime_original" in metadata.tags or tokens.date_stamp:
signals["has_date_stamp"] = True
confidence = 0.95
if "camera_model" in metadata.tags:
signals["camera_model"] = metadata.tags["camera_model"]
confidence = 0.98
# Check if it might be artwork
stem_lower = scanned.path.stem.lower()
if stem_lower in ("cover", "folder", "poster", "fanart", "banner", "front", "back"):
return ClassificationResult(
category="artwork",
confidence=0.95,
signals={"artwork_keyword": stem_lower},
tokens=tokens,
metadata=metadata,
)
return ClassificationResult(
category="photo",
confidence=confidence,
signals=signals,
tokens=tokens,
metadata=metadata,
)
def _classify_audio(
self, scanned: ScannedFile, tokens: TokenizedFilename, metadata: MediaMetadata
) -> ClassificationResult:
path_str = str(scanned.path).lower()
dur = metadata.duration_seconds
tags = metadata.tags
# Check Audiobook indicators
audiobook_score = 0.0
ab_signals = []
if "audiobook" in path_str or "audio books" in path_str:
audiobook_score += 0.4
ab_signals.append("folder_name_audiobook")
if dur > 1800: # > 30 minutes
audiobook_score += 0.3
ab_signals.append("long_duration")
if scanned.path.suffix.lower() == ".m4b":
audiobook_score += 0.5
ab_signals.append("m4b_extension")
if any(k in tags for k in ("narrator", "reader", "series", "composer")):
audiobook_score += 0.2
ab_signals.append("audiobook_tags")
if audiobook_score >= 0.6:
return ClassificationResult(
category="audiobook",
confidence=min(audiobook_score, 0.98),
signals={"audiobook_signals": ab_signals, "duration": dur},
tokens=tokens,
metadata=metadata,
)
# Check Podcast indicators
pod_score = 0.0
pod_signals = []
if "podcast" in path_str or "podcasts" in path_str:
pod_score += 0.4
pod_signals.append("folder_name_podcast")
if (tokens.date_stamp or tokens.air_date) and not tokens.is_photo_or_home_video:
pod_score += 0.45
pod_signals.append("dated_filename")
if tokens.track is None and not tokens.is_music:
pod_score += 0.30
pod_signals.append("non_music_audio_with_date")
elif RE_BROADCAST_DATE.search(scanned.path.stem):
pod_score += 0.45
pod_signals.append("dated_filename")
if tokens.track is None and not tokens.is_music:
pod_score += 0.30
pod_signals.append("non_music_audio_with_date")
if any(k in tags for k in ("podcast", "itunes_category", "show")):
pod_score += 0.35
pod_signals.append("podcast_tags")
if pod_score >= 0.6:
return ClassificationResult(
category="podcast",
confidence=min(pod_score, 0.95),
signals={"podcast_signals": pod_signals},
tokens=tokens,
metadata=metadata,
)
# Standard Music classification
music_score = 0.5 # Base audio file score
m_signals = ["has_audio_stream"]
if tokens.is_music or tokens.track is not None:
music_score += 0.25
m_signals.append("track_number_detected")
if "artist" in tags or tokens.artist:
music_score += 0.15
m_signals.append("artist_present")
if "album" in tags or tokens.album:
music_score += 0.1
m_signals.append("album_present")
if 20 <= dur <= 900: # 20s to 15m typical music track
music_score += 0.1
m_signals.append("typical_song_duration")
if "music" in path_str or "albums" in path_str:
music_score += 0.1
m_signals.append("music_folder_hint")
confidence = min(music_score, 0.99)
needs_quar = confidence < self.confidence_threshold
return ClassificationResult(
category="music",
confidence=confidence,
signals={"music_signals": m_signals, "tags": tags},
tokens=tokens,
metadata=metadata,
needs_quarantine=needs_quar,
quarantine_reason="Audio file lacking track/artist metadata" if needs_quar else None,
)
def _classify_video(
self, scanned: ScannedFile, tokens: TokenizedFilename, metadata: MediaMetadata
) -> ClassificationResult:
path_str = str(scanned.path).lower()
dur = metadata.duration_seconds
stem_lower = scanned.path.stem.lower()
# 1. Home Video Check:
upper_base = scanned.path.stem.split(".")[0].upper()
if upper_base in ("CON", "PRN", "AUX", "NUL"):
return ClassificationResult(
category="home_video",
confidence=0.88,
signals={"reserved_name": upper_base},
tokens=tokens,
metadata=metadata,
)
if (tokens.is_photo_or_home_video or stem_lower.startswith(("vid_", "mov_", "mvi_"))) and (
dur > 0 and dur < 900 or "home" in path_str or "family" in path_str
):
if not tokens.is_episodic and not tokens.resolution:
return ClassificationResult(
category="home_video",
confidence=0.88,
signals={"camera_naming": True, "duration": dur},
tokens=tokens,
metadata=metadata,
)
# 2. Documentary check:
if "documentary" in path_str or "docu" in stem_lower or "bbc." in stem_lower or "national.geographic" in stem_lower:
doc_score = 0.85
if tokens.year:
doc_score += 0.1
return ClassificationResult(
category="documentary",
confidence=min(doc_score, 0.95),
signals={"keyword": "documentary", "year": tokens.year},
tokens=tokens,
metadata=metadata,
)
# 3. Anime Check:
is_standard_tv = bool(RE_STD_TV.search(scanned.path.stem))
is_non_anime_movie = bool(RE_NON_ANIME_GROUPS.search(scanned.path.stem))
has_broadcast_date = bool(tokens.is_daily or tokens.air_date or RE_BROADCAST_DATE.search(scanned.path.stem))
is_anime_candidate = False
a_signals = []
if not is_standard_tv and not is_non_anime_movie and not has_broadcast_date:
if tokens.is_anime:
is_anime_candidate = True
a_signals.append("fansub_syntax")
if RE_ANIME_GROUPS.search(scanned.path.stem) or (tokens.group and tokens.group.lower() in KNOWN_ANIME_GROUPS):
is_anime_candidate = True
a_signals.append("known_anime_group")
if RE_CRC32.search(scanned.path.stem):
is_anime_candidate = True
a_signals.append("crc32_checksum")
if RE_OVA.search(scanned.path.stem):
is_anime_candidate = True
a_signals.append("ova_tag")
if RE_COUR_TAG.search(scanned.path.stem):
is_anime_candidate = True
a_signals.append("cour_tag")
if RE_STANDALONE_EPISODE.search(scanned.path.stem) and not bool(re.search(r"\bseason\b", stem_lower)):
is_anime_candidate = True
a_signals.append("standalone_episode_keyword")
if "anime" in path_str:
is_anime_candidate = True
a_signals.append("anime_folder")
if is_anime_candidate:
anime_score = 0.85
if "known_anime_group" in a_signals:
anime_score += 0.10
if "crc32_checksum" in a_signals:
anime_score += 0.04
conf = min(anime_score, 0.99)
return ClassificationResult(
category="anime",
confidence=conf,
signals={"anime_signals": a_signals},
tokens=tokens,
metadata=metadata,
)
is_tv_folder = bool(re.search(r"(?i)[/\\](?:tv[/\\]|tv[-_\s]shows?|tv[-_\s]series|season[-_\s]*\d+)", path_str))
is_movie_folder = bool(re.search(r"(?i)[/\\](?:movies?[/\\]|films?[/\\])", path_str))
# 4. TV Show Check (Episodic and Daily Broadcasts):
is_daily_tv = has_broadcast_date and (
tokens.resolution
or tokens.source
or "daily" in stem_lower
or "tonight" in stem_lower
or "late" in stem_lower
or "news" in stem_lower
or dur >= 1200
)
is_tv = False
if tokens.is_episodic or is_standard_tv or is_daily_tv:
is_tv = True
elif is_tv_folder and not tokens.year:
is_tv = True
elif is_tv_folder and tokens.episode is not None:
is_tv = True
if is_tv:
tv_score = 0.40
tv_signals = []
if tokens.is_episodic or is_standard_tv:
tv_score += 0.45
tv_signals.append("season_episode_pattern")
if is_daily_tv:
tv_score += 0.45
tv_signals.append("broadcast_date_pattern")
if tokens.episode is not None:
tv_score += 0.10
if is_tv_folder:
tv_score += 0.15
tv_signals.append("tv_folder_hint")
if 600 <= dur <= 5400 and not tokens.year:
tv_score += 0.10
tv_signals.append("episodic_duration")
# Provider boost
prov_res = None
if self.provider and tokens.title and tv_score >= 0.6:
try:
prov_res = self.provider.search_tv(
tokens.title, year=tokens.year, season=tokens.season, episode=tokens.episode
)
if prov_res:
tv_score += prov_res.confidence_boost
tv_signals.append("provider_verified")
except Exception:
pass
conf = min(tv_score, 0.99)
needs_quar = conf < self.confidence_threshold
return ClassificationResult(
category="tv",
confidence=conf,
signals={"tv_signals": tv_signals},
tokens=tokens,
metadata=metadata,
provider_result=prov_res,
needs_quarantine=needs_quar,
quarantine_reason="Low confidence TV classification" if needs_quar else None,
)
# 5. Movie Check:
movie_score = 0.40
m_signals = []
if tokens.year and not has_broadcast_date:
movie_score += 0.35
m_signals.append("year_in_title")
if tokens.resolution or tokens.source or tokens.video_codec:
movie_score += 0.15
m_signals.append("scene_technical_tags")
if is_movie_folder or "movie" in path_str or "film" in path_str:
movie_score += 0.15
m_signals.append("movie_folder_hint")
if dur >= 3600: # > 1 hour
movie_score += 0.20
m_signals.append("feature_film_duration")
# Check for movie extras tag
if re.search(r"-(behindthescenes|deleted|trailer|featurette)\b", stem_lower):
movie_score += 0.25
m_signals.append("movie_extra_tag")
# Provider boost
prov_res = None
if self.provider and tokens.title and movie_score >= 0.5:
try:
prov_res = self.provider.search_movie(tokens.title, year=tokens.year)
if prov_res:
movie_score += prov_res.confidence_boost
m_signals.append("provider_verified")
except Exception:
pass
conf = min(movie_score, 0.99)
needs_quar = conf < self.confidence_threshold
return ClassificationResult(
category="movie",
confidence=conf,
signals={"movie_signals": m_signals, "duration": dur},
tokens=tokens,
metadata=metadata,
provider_result=prov_res,
needs_quarantine=needs_quar,
quarantine_reason="Low confidence movie classification (missing year or title verification)"
if needs_quar
else None,
)
+339
View File
@@ -0,0 +1,339 @@
"""Command-line interface for Media Sorter.
Provides intuitive commands for scanning, dry-run simulation, live atomic organization,
transactional rollback, quarantine resolution, and web dashboard hosting.
"""
from __future__ import annotations
import os
import sys
from pathlib import Path
from typing import Optional
import typer
import uvicorn
from rich.console import Console
from rich.panel import Panel
from rich.progress import BarColumn, Progress, SpinnerColumn, TextColumn, TimeRemainingColumn
from rich.table import Table
from .config import ActionType, Settings
from .db import get_db_session, init_db
from .models import BatchRecord, Operation, QuarantineRecord, QuarantineStatus
from .quarantine import QuarantineManager
from .sorter import MediaSorterApp
app = typer.Typer(
name="media-sorter",
help="Reliable, high-performance media sorter with atomic moves, dry-runs, and rollback.",
add_completion=False,
)
quarantine_app = typer.Typer(help="Manage quarantined or low-confidence review files.")
config_app = typer.Typer(help="Inspect or initialize configuration.")
app.add_typer(quarantine_app, name="quarantine")
app.add_typer(config_app, name="config")
console = Console()
def load_settings_or_default(config_path: Optional[Path] = None) -> Settings:
if config_path and config_path.is_file():
if config_path.suffix in (".env", "") and "env" in config_path.name:
return Settings.load_from_env_file(config_path)
return Settings.load_from_file(config_path)
# Check if .env exists in current working directory
env_file = Path(".env")
if env_file.is_file():
return Settings.load_from_env_file(env_file)
# Search standard configuration paths
for cand in ("config.yaml", "config.yml", "config.toml", "media-sorter.yaml"):
p = Path(cand)
if p.is_file():
return Settings.load_from_file(p)
return Settings()
@app.command()
def scan(
config: Optional[Path] = typer.Option(None, "--config", "-c", help="Path to configuration file"),
source: Optional[Path] = typer.Option(None, "--source", "-s", help="Override source directory"),
):
"""Scan source directories and display media classification preview without moving any files."""
settings = load_settings_or_default(config)
if source:
settings.storage.source_dirs = [str(source)]
engine = init_db(db_path=settings.get_database_path())
sorter = MediaSorterApp(settings, engine)
console.print(Panel(f"[bold cyan]Scanning sources:[/bold cyan] {', '.join(settings.storage.source_dirs)}", title="Media Sorter Discovery"))
with Progress(
SpinnerColumn(),
TextColumn("[progress.description]{task.description}"),
BarColumn(),
TextColumn("[progress.percentage]{task.percentage:>3.0f}%"),
console=console,
) as progress:
task = progress.add_task("[green]Analyzing media...", total=None)
def on_prog(curr, total, name):
progress.update(task, total=total, completed=curr, description=f"[cyan]Probing: {name[:30]}")
results = sorter.scan_and_analyze(progress_callback=on_prog)
if not results:
console.print("[yellow]No qualifying files found in source directories.[/yellow]")
return
table = Table(title=f"Discovered Media Items ({len(results)} total)")
table.add_column("Filename", style="bold", overflow="fold")
table.add_column("Category", style="cyan")
table.add_column("Confidence", justify="right")
table.add_column("Status", justify="center")
for scanned, cls_res in results[:50]:
conf_pct = f"{int(cls_res.confidence * 100)}%"
if cls_res.needs_quarantine:
status_style = "[red]QUARANTINE[/red]"
elif cls_res.confidence >= 0.85:
status_style = "[green]HIGH CONF[/green]"
else:
status_style = "[yellow]MEDIUM[/yellow]"
table.add_row(scanned.path.name, cls_res.category, conf_pct, status_style)
console.print(table)
if len(results) > 50:
console.print(f"[dim]... and {len(results) - 50} more items.[/dim]")
@app.command()
def organize(
config: Optional[Path] = typer.Option(None, "--config", "-c", help="Path to configuration file"),
dry_run: bool = typer.Option(True, "--dry-run/--live", help="Safety preview mode (default: True)"),
source: Optional[Path] = typer.Option(None, "--source", "-s", help="Override source directory"),
dest: Optional[Path] = typer.Option(None, "--dest", "-d", help="Override destination directory"),
action: Optional[ActionType] = typer.Option(None, "--action", "-a", help="Action (move, copy, link, hardlink)"),
threshold: Optional[float] = typer.Option(None, "--threshold", "-t", help="Confidence threshold (0.0-1.0)"),
interval: Optional[int] = typer.Option(None, "--interval", "-i", help="Continuous scan interval in seconds"),
watch: bool = typer.Option(False, "--watch", "-w", help="Run continuously in watch/daemon mode"),
):
"""Execute media organization or generate a dry-run preview."""
import time
settings = load_settings_or_default(config)
if source:
settings.storage.source_dirs = [str(source)]
if dest:
settings.storage.destination_base = str(dest)
if action:
settings.general.action = action
if threshold:
settings.general.confidence_threshold = threshold
loop_interval = interval or (settings.general.scan_interval_seconds if settings.general.scan_interval_seconds > 0 else (60 if watch else 0))
engine = init_db(db_path=settings.get_database_path())
sorter = MediaSorterApp(settings, engine)
mode_label = "[bold yellow]DRY-RUN PREVIEW[/bold yellow]" if dry_run else "[bold red]LIVE EXECUTION[/bold red]"
while True:
console.print(Panel(f"Mode: {mode_label} | Action: {settings.general.action.value.upper()}", title="Media Sorter"))
start_t = time.time()
with Progress(
SpinnerColumn(),
TextColumn("[progress.description]{task.description}"),
BarColumn(),
TextColumn("[progress.percentage]{task.percentage:>3.0f}%"),
TimeRemainingColumn(),
console=console,
) as progress:
task = progress.add_task("[green]Processing...", total=None)
def on_prog(curr, total, name):
progress.update(task, total=total, completed=curr, description=f"Processing: {name[:30]}")
report = sorter.run(dry_run=dry_run, progress_callback=on_prog)
elapsed = max(time.time() - start_t, 0.001)
throughput = round(report.total_files / elapsed, 1)
# Print summary table
summary_table = Table(title=f"Batch Summary [{report.batch_id[:8]}]")
summary_table.add_column("Metric", style="bold")
summary_table.add_column("Count", justify="right")
summary_table.add_row("Total Processed", str(report.total_files))
summary_table.add_row("Moved / Organized", f"[green]{report.moved_files}[/green]")
summary_table.add_row("Copied", f"[blue]{report.copied_files}[/blue]")
summary_table.add_row("Linked", f"[cyan]{report.linked_files}[/cyan]")
summary_table.add_row("Skipped (Conflicts / Existing)", f"[dim]{report.skipped_files}[/dim]")
summary_table.add_row("Quarantined (Review Queue)", f"[yellow]{report.quarantined_files}[/yellow]")
summary_table.add_row("Failures", f"[red]{report.failed_files}[/red]")
summary_table.add_row("Throughput", f"{throughput} files/sec ({round(elapsed, 2)}s)")
console.print(summary_table)
if dry_run:
console.print("\n[bold cyan]Safe dry-run complete. No files were modified on disk.[/bold cyan]")
console.print("[dim]To apply these changes live, rerun with --live.[/dim]")
if loop_interval <= 0:
break
console.print(f"\n[cyan]Sleeping for {loop_interval}s until next scan cycle (press Ctrl+C to stop)...[/cyan]")
try:
time.sleep(loop_interval)
except KeyboardInterrupt:
console.print("\n[yellow]Daemon watch loop stopped by user.[/yellow]")
break
@app.command()
def rollback(
batch_id: Optional[str] = typer.Option(None, "--batch-id", "-b", help="Specific batch ID to roll back"),
config: Optional[Path] = typer.Option(None, "--config", "-c", help="Path to configuration file"),
):
"""Roll back a previous organization batch, restoring moved files to original sources."""
settings = load_settings_or_default(config)
engine = init_db(db_path=settings.get_database_path())
sorter = MediaSorterApp(settings, engine)
with console.status("[bold yellow]Executing transactional rollback...[/bold yellow]"):
reverted = sorter.rollback(batch_id)
if reverted > 0:
console.print(f"[bold green]Successfully rolled back {reverted} file operations.[/bold green]")
else:
console.print("[yellow]No operations were reverted (batch already rolled back or not found).[/yellow]")
@app.command()
def history(
limit: int = typer.Option(10, "--limit", "-n", help="Number of past batches to display"),
config: Optional[Path] = typer.Option(None, "--config", "-c", help="Path to configuration file"),
):
"""Display history of past organization batches and execution logs."""
settings = load_settings_or_default(config)
engine = init_db(db_path=settings.get_database_path())
with get_db_session(engine) as session:
batches = session.query(BatchRecord).order_by(BatchRecord.created_at.desc()).limit(limit).all()
if not batches:
console.print("[dim]No batch records found.[/dim]")
return
table = Table(title=f"Execution History (Last {len(batches)})")
table.add_column("Batch ID", style="bold")
table.add_column("Date", style="dim")
table.add_column("Mode")
table.add_column("Status")
table.add_column("Total", justify="right")
table.add_column("Moved", justify="right")
table.add_column("Quarantined", justify="right")
for b in batches:
mode = "[dim]Dry-Run[/dim]" if b.dry_run else "[bold]Live[/bold]"
status = f"[green]{b.status}[/green]" if b.status == "COMPLETED" else f"[yellow]{b.status}[/yellow]"
date_str = b.created_at.strftime("%Y-%m-%d %H:%M") if b.created_at else "-"
table.add_row(b.id[:8], date_str, mode, status, str(b.total_files), str(b.moved_files), str(b.quarantined_files))
console.print(table)
@quarantine_app.command("list")
def quarantine_list(
config: Optional[Path] = typer.Option(None, "--config", "-c", help="Path to configuration file"),
):
"""List pending items requiring manual review."""
settings = load_settings_or_default(config)
engine = init_db(db_path=settings.get_database_path())
with get_db_session(engine) as session:
qm = QuarantineManager(session)
items = qm.list_pending()
if not items:
console.print("[green]Quarantine queue is empty. All media classified cleanly![/green]")
return
table = Table(title=f"Quarantine Review Queue ({len(items)} items)")
table.add_column("ID", justify="right")
table.add_column("File Path", overflow="fold")
table.add_column("Suggested", style="cyan")
table.add_column("Confidence", justify="right")
table.add_column("Reason", style="yellow")
for q in items:
conf = f"{int((q.confidence or 0) * 100)}%"
table.add_row(str(q.id), q.src, q.suggested_category or "unknown", conf, q.reason)
console.print(table)
@quarantine_app.command("resolve")
def quarantine_resolve(
item_id: int = typer.Argument(..., help="Quarantine record ID to resolve"),
category: str = typer.Option(..., "--category", "-cat", help="Target category (movie, tv, music, etc.)"),
config: Optional[Path] = typer.Option(None, "--config", "-c", help="Path to configuration file"),
):
"""Manually classify and resolve a quarantined item."""
settings = load_settings_or_default(config)
engine = init_db(db_path=settings.get_database_path())
with get_db_session(engine) as session:
qm = QuarantineManager(session)
success = qm.resolve_item(item_id, category)
if success:
console.print(f"[green]Successfully resolved item #{item_id} as {category}.[/green]")
else:
console.print(f"[red]Quarantine item #{item_id} not found.[/red]")
@app.command()
def server(
host: Optional[str] = typer.Option(None, "--host", "-h", help="Bind host (default from .env or 0.0.0.0)"),
port: Optional[int] = typer.Option(None, "--port", "-p", help="Bind port (default from .env or 8080)"),
config: Optional[Path] = typer.Option(None, "--config", "-c", help="Path to configuration file"),
):
"""Launch web dashboard and management server."""
settings = load_settings_or_default(config)
from .server import create_app
bind_host = host or settings.server.host or "0.0.0.0"
bind_port = port or settings.server.port or 8080
web_app = create_app(settings)
console.print(f"[bold green]Starting Media Sorter Dashboard on http://{bind_host}:{bind_port}[/bold green]")
uvicorn.run(web_app, host=bind_host, port=bind_port)
@config_app.command("show")
def config_show(
config: Optional[Path] = typer.Option(None, "--config", "-c", help="Path to configuration file"),
):
"""Print effective configuration settings."""
settings = load_settings_or_default(config)
import yaml
console.print(yaml.dump(settings.model_dump(mode="json"), default_flow_style=False, sort_keys=False))
@config_app.command("init")
def config_init(
output: Path = typer.Option(Path("media-sorter.yaml"), "--output", "-o", help="Target config file path"),
):
"""Create a safe starter configuration file."""
if output.exists():
console.print(f"[yellow]Configuration file already exists at {output}. Aborting.[/yellow]")
return
settings = Settings()
settings.dump_yaml(output)
console.print(f"[green]Created default configuration file at {output}.[/green]")
if __name__ == "__main__":
app()
+497
View File
@@ -0,0 +1,497 @@
"""Configuration management for Media Sorter.
Provides Pydantic-based settings validated from YAML, TOML, JSON, or environment variables.
"""
from __future__ import annotations
import os
import json
from enum import Enum
from pathlib import Path
from typing import Any, Dict, List, Optional
import yaml
from pydantic import BaseModel, Field, field_validator
from pydantic_settings import BaseSettings, SettingsConfigDict
class ActionType(str, Enum):
MOVE = "move"
COPY = "copy"
LINK = "link"
HARDLINK = "hardlink"
class ConflictPolicy(str, Enum):
SKIP = "skip"
RENAME_UNIQUE = "rename_unique"
QUARANTINE = "quarantine"
REPLACE_IF_HIGHER_QUALITY = "replace_if_higher_quality"
ERROR = "error"
class DestinationDirs(BaseModel):
movies: str = "Movies"
tv: str = "TV Shows"
anime: str = "Anime"
music: str = "Music"
audiobooks: str = "Audiobooks"
podcasts: str = "Podcasts"
home_videos: str = "Home Videos"
photos: str = "Photos"
archives: str = "Archives"
quarantine: str = "Quarantine"
class GeneralSettings(BaseModel):
dry_run: bool = True
confidence_threshold: float = 0.75
worker_count: int = 4
min_file_age_seconds: int = 300
action: ActionType = ActionType.MOVE
preserve_permissions: bool = True
log_level: str = "INFO"
scan_interval_seconds: int = 0
cleanup_empty_dirs: bool = True
rename_files: bool = True
@field_validator("confidence_threshold")
@classmethod
def validate_confidence(cls, v: float) -> float:
if not 0.0 < v <= 1.0:
raise ValueError("confidence_threshold must be between 0.0 and 1.0")
return v
@field_validator("worker_count")
@classmethod
def validate_worker_count(cls, v: int) -> int:
if v < 1:
raise ValueError("worker_count must be at least 1")
return v
@field_validator("min_file_age_seconds")
@classmethod
def validate_min_file_age(cls, v: int) -> int:
if v < 0:
raise ValueError("min_file_age_seconds cannot be negative")
return v
@field_validator("log_level")
@classmethod
def validate_log_level(cls, v: str) -> str:
allowed = {"DEBUG", "INFO", "WARNING", "ERROR", "CRITICAL"}
upper = v.upper()
if upper not in allowed:
raise ValueError(f"log_level must be one of {allowed}")
return upper
class StorageSettings(BaseModel):
source_dirs: List[str] = Field(default_factory=lambda: ["incoming"])
destination_base: str = "organized"
destination_dirs: DestinationDirs = Field(default_factory=DestinationDirs)
class ConflictSettings(BaseModel):
policy: ConflictPolicy = ConflictPolicy.RENAME_UNIQUE
allow_overwrite: bool = False
backup_dir: Optional[str] = None
class FilterSettings(BaseModel):
include_patterns: List[str] = Field(default_factory=lambda: ["*"])
exclude_patterns: List[str] = Field(
default_factory=lambda: [
".*",
"*.part",
"*.crdownload",
"*.!qB",
"Thumbs.db",
"desktop.ini",
"@eaDir",
"$RECYCLE.BIN",
"*.txt",
]
)
class TemplateSettings(BaseModel):
movie: str = "{title} ({year})/{movie_name}.{ext}"
tv: str = "{title}/Season {season:02d}/{show_name}_{season_episode}.{ext}"
anime: str = "{title}/Season {season:02d}/{show_name}_{season_episode} [{group}].{ext}"
music: str = "{artist}/{album} ({year})/{disc:01d}{track:02d} - {title}.{ext}"
audiobook: str = "{author}/{title}/{track:02d} - {chapter}.{ext}"
podcast: str = "{show}/{year}/{show} - {date} - {title}.{ext}"
home_video: str = "{year}/{year}-{month:02d} - {event}/{filename}.{ext}"
photo: str = "{year}/{year}-{month:02d}/{year}{month:02d}{day:02d}_{time}_{camera}.{ext}"
archive: str = "Archives/{filename}.{ext}"
quarantine: str = "Quarantine/{reason}/{filename}.{ext}"
class SubtitleSettings(BaseModel):
match_video_basename: bool = True
preserve_language_code: bool = True
class ArtworkSettings(BaseModel):
match_parent_folder: bool = True
class ExtrasSettings(BaseModel):
detect_trailers: bool = True
trailer_suffix: str = "-trailer"
class SidecarSettings(BaseModel):
enabled: bool = True
subtitles: SubtitleSettings = Field(default_factory=SubtitleSettings)
artwork: ArtworkSettings = Field(default_factory=ArtworkSettings)
extras: ExtrasSettings = Field(default_factory=ExtrasSettings)
class DatabaseSettings(BaseModel):
path: str = "media_sorter.db"
wal_mode: bool = True
class ServerSettings(BaseModel):
host: str = "127.0.0.1"
port: int = 8080
enabled: bool = True
class ProviderSettings(BaseModel):
enable_online_metadata: bool = False
tmdb_api_key: Optional[str] = None
tvdb_api_key: Optional[str] = None
rate_limit_per_second: float = 2.0
cache_expiry_hours: int = 72
class NotificationSettings(BaseModel):
enabled: bool = False
webhook_url: Optional[str] = None
notify_on_complete: bool = True
notify_on_failure: bool = True
class SymlinkSettings(BaseModel):
follow_symlinks: bool = False
handle_broken_symlinks: str = "skip" # skip | quarantine
class PermissionSettings(BaseModel):
preserve_attributes: bool = True
file_mode: Optional[str] = None # e.g. "0644"
dir_mode: Optional[str] = None # e.g. "0755"
owner: Optional[str] = None
group: Optional[str] = None
class QuarantineSettings(BaseModel):
move_to_quarantine_folder: bool = False
directory: str = "Quarantine"
class Settings(BaseSettings):
model_config = SettingsConfigDict(
env_prefix="MEDIA_SORTER_",
env_nested_delimiter="__",
env_file=".env",
env_file_encoding="utf-8",
extra="ignore",
)
general: GeneralSettings = Field(default_factory=GeneralSettings)
storage: StorageSettings = Field(default_factory=StorageSettings)
conflicts: ConflictSettings = Field(default_factory=ConflictSettings)
filters: FilterSettings = Field(default_factory=FilterSettings)
templates: TemplateSettings = Field(default_factory=TemplateSettings)
sidecars: SidecarSettings = Field(default_factory=SidecarSettings)
database: DatabaseSettings = Field(default_factory=DatabaseSettings)
server: ServerSettings = Field(default_factory=ServerSettings)
providers: ProviderSettings = Field(default_factory=ProviderSettings)
notifications: NotificationSettings = Field(default_factory=NotificationSettings)
symlinks: SymlinkSettings = Field(default_factory=SymlinkSettings)
permissions: PermissionSettings = Field(default_factory=PermissionSettings)
quarantine: QuarantineSettings = Field(default_factory=QuarantineSettings)
def model_post_init(self, __context: Any) -> None:
super().model_post_init(__context)
# Check intuitive environment variable overrides from .env only if not explicitly supplied
if "storage" not in self.model_fields_set:
downloads_dir = os.getenv("DOWNLOADS_DIR") or os.getenv("SOURCE_DIR")
if downloads_dir:
self.storage.source_dirs = [downloads_dir]
movies_dir = os.getenv("MOVIES_DIR")
if movies_dir:
self.storage.destination_dirs.movies = movies_dir
shows_dir = os.getenv("SHOWS_DIR") or os.getenv("TV_DIR")
if shows_dir:
self.storage.destination_dirs.tv = shows_dir
anime_dir = os.getenv("ANIME_DIR")
if anime_dir:
self.storage.destination_dirs.anime = anime_dir
elif shows_dir:
self.storage.destination_dirs.anime = shows_dir
if "general" not in self.model_fields_set:
dry_run_env = os.getenv("DRY_RUN")
if dry_run_env is not None:
self.general.dry_run = dry_run_env.strip().lower() in ("true", "1", "yes", "on")
action_env = os.getenv("ACTION")
if action_env:
try:
self.general.action = ActionType(action_env.lower())
except ValueError:
pass
conf_env = os.getenv("CONFIDENCE_THRESHOLD")
if conf_env:
try:
self.general.confidence_threshold = float(conf_env)
except ValueError:
pass
min_age_env = os.getenv("MIN_FILE_AGE_SECONDS")
if min_age_env:
try:
self.general.min_file_age_seconds = int(min_age_env)
except ValueError:
pass
scan_int_env = os.getenv("SCAN_INTERVAL_SECONDS")
if scan_int_env:
try:
self.general.scan_interval_seconds = int(scan_int_env)
except ValueError:
pass
cleanup_env = os.getenv("CLEANUP_EMPTY_DIRS")
if cleanup_env is not None:
self.general.cleanup_empty_dirs = cleanup_env.strip().lower() in ("true", "1", "yes", "on")
rename_env = os.getenv("RENAME_FILES")
if rename_env is not None:
self.general.rename_files = rename_env.strip().lower() in ("true", "1", "yes", "on")
if "templates" not in self.model_fields_set:
movie_tmpl = os.getenv("MOVIE_TEMPLATE")
if movie_tmpl:
self.templates.movie = movie_tmpl
tv_tmpl = os.getenv("TV_TEMPLATE") or os.getenv("SHOW_TEMPLATE") or os.getenv("SHOWS_TEMPLATE")
if tv_tmpl:
self.templates.tv = tv_tmpl
if "server" not in self.model_fields_set:
host_env = os.getenv("SERVER_HOST") or os.getenv("HOST")
if host_env:
self.server.host = host_env
port_env = os.getenv("SERVER_PORT") or os.getenv("PORT")
if port_env:
try:
self.server.port = int(port_env)
except ValueError:
pass
if "database" not in self.model_fields_set:
db_path_env = os.getenv("DATABASE_PATH")
if db_path_env:
self.database.path = db_path_env
def resolve_path(self, raw_path: str) -> Path:
"""Expand environment variables and user home, returning resolved Path."""
expanded = os.path.expandvars(raw_path)
return Path(expanded).expanduser().resolve()
def get_source_paths(self) -> List[Path]:
return [self.resolve_path(p) for p in self.storage.source_dirs]
def get_destination_base_path(self) -> Path:
return self.resolve_path(self.storage.destination_base)
def get_destination_path(self, category: str) -> Path:
"""Return the destination path for a given category."""
base = self.get_destination_base_path()
cat_map = {
"movie": "movies",
"movies": "movies",
"tv": "tv",
"show": "tv",
"shows": "tv",
"audiobook": "audiobooks",
"podcast": "podcasts",
"photo": "photos",
"home_video": "home_videos",
"archive": "archives",
}
lookup_key = cat_map.get(category, category)
dest_field = getattr(self.storage.destination_dirs, lookup_key, category)
path = Path(dest_field)
# If dest_field is an explicit relative path (e.g. ./movies, ./shows) or absolute path
if path.is_absolute() or str(dest_field).startswith(("./", "../")):
return path.resolve()
return (base / path).resolve()
@classmethod
def load_from_env_file(cls, env_path: Path | str = ".env") -> Settings:
"""Load configuration from a .env file."""
path = Path(env_path).expanduser().resolve()
settings = cls()
if not path.is_file():
return settings
from dotenv import dotenv_values
values = dotenv_values(path)
downloads_dir = os.getenv("DOWNLOADS_DIR") or values.get("DOWNLOADS_DIR") or os.getenv("SOURCE_DIR") or values.get("SOURCE_DIR")
if downloads_dir:
settings.storage.source_dirs = [downloads_dir]
movies_dir = os.getenv("MOVIES_DIR") or values.get("MOVIES_DIR")
if movies_dir:
settings.storage.destination_dirs.movies = movies_dir
shows_dir = os.getenv("SHOWS_DIR") or values.get("SHOWS_DIR") or os.getenv("TV_DIR") or values.get("TV_DIR")
if shows_dir:
settings.storage.destination_dirs.tv = shows_dir
anime_dir = os.getenv("ANIME_DIR") or values.get("ANIME_DIR")
if anime_dir:
settings.storage.destination_dirs.anime = anime_dir
elif shows_dir:
settings.storage.destination_dirs.anime = shows_dir
dry_run = os.getenv("DRY_RUN") if os.getenv("DRY_RUN") is not None else values.get("DRY_RUN")
if dry_run is not None:
settings.general.dry_run = dry_run.strip().lower() in ("true", "1", "yes", "on")
action = os.getenv("ACTION") or values.get("ACTION")
if action:
try:
settings.general.action = ActionType(action.lower())
except ValueError:
pass
conf = os.getenv("CONFIDENCE_THRESHOLD") or values.get("CONFIDENCE_THRESHOLD")
if conf:
try:
settings.general.confidence_threshold = float(conf)
except ValueError:
pass
min_age = os.getenv("MIN_FILE_AGE_SECONDS") or values.get("MIN_FILE_AGE_SECONDS")
if min_age:
try:
settings.general.min_file_age_seconds = int(min_age)
except ValueError:
pass
scan_int = os.getenv("SCAN_INTERVAL_SECONDS") or values.get("SCAN_INTERVAL_SECONDS")
if scan_int:
try:
settings.general.scan_interval_seconds = int(scan_int)
except ValueError:
pass
cleanup = os.getenv("CLEANUP_EMPTY_DIRS") if os.getenv("CLEANUP_EMPTY_DIRS") is not None else values.get("CLEANUP_EMPTY_DIRS")
if cleanup is not None:
settings.general.cleanup_empty_dirs = str(cleanup).strip().lower() in ("true", "1", "yes", "on")
rename_files = os.getenv("RENAME_FILES") if os.getenv("RENAME_FILES") is not None else values.get("RENAME_FILES")
if rename_files is not None:
settings.general.rename_files = str(rename_files).strip().lower() in ("true", "1", "yes", "on")
movie_tmpl = os.getenv("MOVIE_TEMPLATE") or values.get("MOVIE_TEMPLATE")
if movie_tmpl:
settings.templates.movie = movie_tmpl
tv_tmpl = os.getenv("TV_TEMPLATE") or values.get("TV_TEMPLATE") or os.getenv("SHOW_TEMPLATE") or values.get("SHOW_TEMPLATE") or os.getenv("SHOWS_TEMPLATE") or values.get("SHOWS_TEMPLATE")
if tv_tmpl:
settings.templates.tv = tv_tmpl
host = os.getenv("SERVER_HOST") or values.get("SERVER_HOST") or os.getenv("HOST") or values.get("HOST")
if host:
settings.server.host = host
port = os.getenv("SERVER_PORT") or values.get("SERVER_PORT") or os.getenv("PORT") or values.get("PORT")
if port:
try:
settings.server.port = int(port)
except ValueError:
pass
db_path = os.getenv("DATABASE_PATH") or values.get("DATABASE_PATH")
if db_path:
settings.database.path = db_path
return settings
def save_to_env_file(self, env_path: Path | str = ".env") -> None:
"""Persist key user-configurable settings to a .env file."""
path = Path(env_path)
content = (
f"# Media Sorter Configuration\n"
f"DOWNLOADS_DIR={self.storage.source_dirs[0] if self.storage.source_dirs else './downloads'}\n"
f"MOVIES_DIR={self.storage.destination_dirs.movies}\n"
f"SHOWS_DIR={self.storage.destination_dirs.tv}\n"
f"DRY_RUN={'true' if self.general.dry_run else 'false'}\n"
f"ACTION={self.general.action.value}\n"
f"CONFIDENCE_THRESHOLD={self.general.confidence_threshold}\n"
f"MIN_FILE_AGE_SECONDS={self.general.min_file_age_seconds}\n"
f"SCAN_INTERVAL_SECONDS={self.general.scan_interval_seconds}\n"
f"CLEANUP_EMPTY_DIRS={'true' if self.general.cleanup_empty_dirs else 'false'}\n"
f"RENAME_FILES={'true' if self.general.rename_files else 'false'}\n"
f"MOVIE_TEMPLATE={self.templates.movie}\n"
f"TV_TEMPLATE={self.templates.tv}\n"
f"SERVER_HOST={self.server.host}\n"
f"SERVER_PORT={self.server.port}\n"
f"DATABASE_PATH={self.database.path}\n"
)
path.write_text(content, encoding="utf-8")
def get_database_path(self) -> Path:
return self.resolve_path(self.database.path)
@classmethod
def load_from_file(cls, config_path: Path | str) -> Settings:
"""Load configuration from YAML, TOML, or JSON file."""
path = Path(config_path).expanduser().resolve()
if not path.is_file():
raise FileNotFoundError(f"Configuration file not found: {path}")
ext = path.suffix.lower()
with open(path, "r", encoding="utf-8") as f:
content = f.read()
if ext in (".yaml", ".yml"):
data = yaml.safe_load(content) or {}
elif ext == ".json":
data = json.loads(content)
elif ext == ".toml":
try:
import tomllib # Python 3.11+
data = tomllib.loads(content)
except ImportError:
import tomli
data = tomli.loads(content)
else:
# Fallback to YAML loader which can parse JSON and YAML
data = yaml.safe_load(content) or {}
return cls(**data)
def dump_yaml(self, target_path: Path | str) -> None:
"""Dump settings to YAML format."""
path = Path(target_path)
path.parent.mkdir(parents=True, exist_ok=True)
data = self.model_dump(mode="json")
with open(path, "w", encoding="utf-8") as f:
yaml.dump(data, f, default_flow_style=False, sort_keys=False)
+84
View File
@@ -0,0 +1,84 @@
"""Database initialization and session management for Media Sorter.
Configures SQLite with Write-Ahead Logging (WAL) mode and foreign keys enabled
for transactional safety, high concurrency, and crash resilience.
"""
from __future__ import annotations
import os
from contextlib import contextmanager
from pathlib import Path
from typing import Generator, Optional
from sqlalchemy import create_engine, event
from sqlalchemy.engine import Engine
from sqlalchemy.orm import Session, declarative_base, scoped_session, sessionmaker
from .models import Base
DEFAULT_DB_PATH = Path("media_sorter.db")
def get_engine(db_path: Path | str = DEFAULT_DB_PATH, wal_mode: bool = True) -> Engine:
"""Create and configure a SQLite SQLAlchemy engine.
Enables WAL mode and enforces foreign keys for transactional integrity.
"""
path = Path(db_path).resolve()
path.parent.mkdir(parents=True, exist_ok=True)
db_url = f"sqlite:///{path}"
engine = create_engine(
db_url,
connect_args={"check_same_thread": False, "timeout": 30.0},
pool_pre_ping=True,
)
@event.listens_for(engine, "connect")
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
if wal_mode:
cursor.execute("PRAGMA journal_mode=WAL")
cursor.execute("PRAGMA synchronous=NORMAL")
cursor.execute("PRAGMA busy_timeout=30000")
cursor.close()
return engine
def init_db(engine: Optional[Engine] = None, db_path: Path | str = DEFAULT_DB_PATH) -> Engine:
"""Initialize all tables defined in models.py if they do not exist."""
if engine is None:
engine = get_engine(db_path)
Base.metadata.create_all(engine)
return engine
_ENGINE_SESSION_FACTORIES: dict[Engine, scoped_session[Session]] = {}
def get_session_factory(engine: Engine) -> scoped_session[Session]:
"""Retrieve or create a cached scoped session factory bound to the given engine."""
if engine not in _ENGINE_SESSION_FACTORIES:
_ENGINE_SESSION_FACTORIES[engine] = scoped_session(
sessionmaker(autocommit=False, autoflush=False, bind=engine)
)
return _ENGINE_SESSION_FACTORIES[engine]
@contextmanager
def get_db_session(engine: Engine) -> Generator[Session, None, None]:
"""Provide a transactional scope around a series of operations."""
session_factory = get_session_factory(engine)
session: Session = session_factory()
try:
yield session
session.commit()
except Exception:
session.rollback()
raise
finally:
session.close()
session_factory.remove()
+791
View File
@@ -0,0 +1,791 @@
"""Safe execution engine and transactional operation journal for Media Sorter.
Enforces dry-run previews, atomic moves, cross-filesystem safety, conflict handling,
attribute preservation (POSIX timestamps/permissions), crash recovery, and instant rollbacks.
"""
from __future__ import annotations
import hashlib
import os
import shutil
import uuid
from dataclasses import dataclass, field
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Callable, Dict, List, Optional
import structlog
from sqlalchemy.orm import Session
from .config import ActionType, ConflictPolicy, Settings
from .models import BatchRecord, FileRecord, Operation, OperationStatus, QuarantineRecord, QuarantineStatus
from contextlib import contextmanager
import sys
logger = structlog.get_logger(__name__)
class ProcessLockError(Exception):
"""Raised when another media-sorter process holds the execution lock."""
pass
@contextmanager
def acquire_process_lock(lock_file_path: Path):
"""Ensure mutual exclusion so multiple workers/instances do not run concurrent batches."""
lock_file = Path(lock_file_path).resolve()
lock_file.parent.mkdir(parents=True, exist_ok=True)
f = open(lock_file, "a+")
try:
if sys.platform != "win32":
import fcntl
try:
fcntl.flock(f.fileno(), fcntl.LOCK_EX | fcntl.LOCK_NB)
except (IOError, OSError):
raise ProcessLockError(
f"Another media-sorter process currently holds the lock on {lock_file}."
)
else:
import msvcrt
try:
msvcrt.locking(f.fileno(), msvcrt.LK_NBLCK, 1)
except (IOError, OSError):
raise ProcessLockError(
f"Another media-sorter process currently holds the lock on {lock_file}."
)
yield
finally:
try:
if sys.platform != "win32":
import fcntl
fcntl.flock(f.fileno(), fcntl.LOCK_UN)
else:
import msvcrt
msvcrt.locking(f.fileno(), msvcrt.LK_UNLCK, 1)
except Exception:
pass
f.close()
def copy_extended_attributes(src: Path, dst: Path) -> None:
"""Preserve POSIX extended attributes (xattrs) across filesystems where supported."""
if hasattr(os, "listxattr") and hasattr(os, "getxattr") and hasattr(os, "setxattr"):
try:
attrs = os.listxattr(src)
for attr in attrs:
try:
val = os.getxattr(src, attr)
os.setxattr(dst, attr, val)
except (OSError, PermissionError):
pass
except (OSError, PermissionError):
pass
def compute_file_hash(file_path: Path, max_bytes: int = 1048576) -> str:
"""Compute quick partial SHA-256 hash (first 1MB) for fast identity verification."""
try:
hasher = hashlib.sha256()
with open(file_path, "rb") as f:
chunk = f.read(max_bytes)
hasher.update(chunk)
return hasher.hexdigest()
except Exception:
return ""
@dataclass
class PlannedOperation:
src: Path
dst: Path
action: ActionType
category: str
confidence: float
details: Dict[str, Any] = field(default_factory=dict)
is_conflict: bool = False
conflict_resolved_dst: Optional[Path] = None
quarantine: bool = False
quarantine_reason: Optional[str] = None
primary_src: Optional[Path] = None
@dataclass
class BatchExecutionReport:
batch_id: str
dry_run: bool
total_files: int = 0
moved_files: int = 0
copied_files: int = 0
linked_files: int = 0
skipped_files: int = 0
quarantined_files: int = 0
failed_files: int = 0
cleaned_dirs: int = 0
operations: List[PlannedOperation] = field(default_factory=list)
errors: List[str] = field(default_factory=list)
class MediaExecutor:
"""Executes planned operations with transactional safety and crash recovery."""
def __init__(self, settings: Settings, session: Session):
self.settings = settings
self.session = session
def plan_operations(
self, planned_items: List[PlannedOperation]
) -> List[PlannedOperation]:
"""Validate destination conflicts and resolve destination paths."""
allocated_destinations: Dict[Path, PlannedOperation] = {}
validated_plan: List[PlannedOperation] = []
primary_resolutions: Dict[Path, Path] = {}
primary_orig_to_resolved: Dict[Path, Path] = {}
for item in planned_items:
# If already marked for quarantine, keep as is
if item.quarantine:
validated_plan.append(item)
continue
# Check if this item is a sidecar whose primary was renamed
orig_target_dst = item.dst
if item.category in ("subtitle", "artwork", "metadata") or item.primary_src:
p_src = item.primary_src
resolved_p_dst = None
if p_src and p_src in primary_resolutions:
resolved_p_dst = primary_resolutions[p_src]
else:
for orig_p_dst, res_p_dst in primary_orig_to_resolved.items():
if orig_p_dst.parent == item.dst.parent and item.dst.stem.lower().startswith(orig_p_dst.stem.lower()):
resolved_p_dst = res_p_dst
break
if resolved_p_dst and resolved_p_dst.stem != item.dst.stem:
orig_stem = item.dst.stem
p_orig_stem = None
for orig_p_dst in primary_orig_to_resolved:
if orig_stem.lower().startswith(orig_p_dst.stem.lower()):
p_orig_stem = orig_p_dst.stem
break
tag = orig_stem[len(p_orig_stem):] if p_orig_stem else ""
new_sidecar_name = f"{resolved_p_dst.stem}{tag}{item.dst.suffix}"
item.dst = resolved_p_dst.parent / new_sidecar_name
target_dst = item.dst
# 1. Check intra-batch duplicate destination conflict
if target_dst in allocated_destinations:
logger.warning(
"Intra-batch destination collision detected",
dst=str(target_dst),
src1=str(allocated_destinations[target_dst].src),
src2=str(item.src),
)
item.is_conflict = True
target_dst = self._resolve_conflict(item.src, target_dst)
item.conflict_resolved_dst = target_dst
# 2. Check on-disk destination conflict
if target_dst.exists():
logger.info("Destination already exists on disk", dst=str(target_dst), src=str(item.src))
item.is_conflict = True
target_dst = self._resolve_conflict(item.src, target_dst)
item.conflict_resolved_dst = target_dst
if item.quarantine:
validated_plan.append(item)
continue
item.dst = target_dst
allocated_destinations[target_dst] = item
validated_plan.append(item)
# Record primary resolution for companion alignment
if item.category not in ("subtitle", "artwork", "metadata"):
primary_resolutions[item.src] = item.dst
primary_orig_to_resolved[orig_target_dst] = item.dst
return validated_plan
def _resolve_conflict(self, src: Path, desired_dst: Path) -> Path:
"""Resolve conflict according to the configured conflict policy."""
policy = self.settings.conflicts.policy
if policy == ConflictPolicy.SKIP:
return desired_dst # Will be skipped during execution
elif policy == ConflictPolicy.ERROR:
raise FileExistsError(f"Destination conflict: {desired_dst} already exists")
elif policy == ConflictPolicy.QUARANTINE:
return self.settings.get_destination_path("quarantine") / f"conflicts/{src.name}"
elif policy == ConflictPolicy.REPLACE_IF_HIGHER_QUALITY:
# Allow replacing existing file (will backup during execution)
return desired_dst
else:
# ConflictPolicy.RENAME_UNIQUE: foo (1).mp4
parent = desired_dst.parent
stem = desired_dst.stem
ext = desired_dst.suffix
counter = 1
candidate = parent / f"{stem} ({counter}){ext}"
while candidate.exists():
counter += 1
candidate = parent / f"{stem} ({counter}){ext}"
return candidate
def execute_batch(
self,
planned_items: List[PlannedOperation],
dry_run: Optional[bool] = None,
progress_callback: Optional[Callable[[int, int, str], None]] = None,
) -> BatchExecutionReport:
"""Execute a batch of operations transactionally, with dry-run support."""
is_dry_run = self.settings.general.dry_run if dry_run is None else dry_run
batch_id = str(uuid.uuid4())
# Validate and resolve destination collisions
validated_plan = self.plan_operations(planned_items)
report = BatchExecutionReport(
batch_id=batch_id,
dry_run=is_dry_run,
total_files=len(validated_plan),
operations=validated_plan,
)
# Create Batch Record in DB
batch_record = BatchRecord(
id=batch_id,
dry_run=is_dry_run,
status="IN_PROGRESS",
total_files=len(validated_plan),
)
self.session.add(batch_record)
self.session.commit()
total = len(validated_plan)
for idx, item in enumerate(validated_plan):
if progress_callback:
progress_callback(idx + 1, total, str(item.src.name))
if item.quarantine:
self._record_quarantine(batch_id, item, is_dry_run)
report.quarantined_files += 1
continue
# Check if skipping due to conflict
if item.is_conflict and self.settings.conflicts.policy == ConflictPolicy.SKIP and item.dst.exists():
logger.info("Skipping existing destination", dst=str(item.dst))
report.skipped_files += 1
self._record_operation(
batch_id, item, status=OperationStatus.SKIPPED, is_dry_run=is_dry_run
)
continue
if is_dry_run:
# Dry run preview only: do not touch filesystem
if item.action == ActionType.MOVE:
report.moved_files += 1
elif item.action == ActionType.COPY:
report.copied_files += 1
elif item.action in (ActionType.LINK, ActionType.HARDLINK):
report.linked_files += 1
self._record_operation(
batch_id, item, status=OperationStatus.PLANNED, is_dry_run=True
)
continue
# Live execution
try:
self._execute_single_op(batch_id, item)
if item.action == ActionType.MOVE:
report.moved_files += 1
elif item.action == ActionType.COPY:
report.copied_files += 1
elif item.action in (ActionType.LINK, ActionType.HARDLINK):
report.linked_files += 1
except Exception as e:
report.failed_files += 1
err_msg = f"Failed {item.action} on {item.src} -> {item.dst}: {e}"
logger.error(err_msg, exc_info=True)
report.errors.append(err_msg)
# Clean up empty directories in source directories after live moves
if not is_dry_run and getattr(self.settings.general, "cleanup_empty_dirs", True):
moved_srcs = [
item.src
for item in validated_plan
if item.action == ActionType.MOVE and not item.quarantine
]
report.cleaned_dirs = self.clean_empty_directories(moved_srcs)
# Update batch record completion status
batch_record.completed_at = datetime.now(timezone.utc)
batch_record.moved_files = report.moved_files
batch_record.skipped_files = report.skipped_files
batch_record.failed_files = report.failed_files
batch_record.quarantined_files = report.quarantined_files
batch_record.status = "COMPLETED" if report.failed_files == 0 else "PARTIAL_FAILURE"
self.session.commit()
return report
def _execute_single_op(self, batch_id: str, item: PlannedOperation) -> None:
"""Perform atomic move, copy, or link with attribute preservation and journal update."""
src = item.src
dst = item.dst
action = item.action
if not src.exists():
raise FileNotFoundError(f"Source file missing: {src}")
src_hash = compute_file_hash(src)
backup_path: Optional[str] = None
# Handle backup if replacing
if dst.exists():
if self.settings.conflicts.policy == ConflictPolicy.REPLACE_IF_HIGHER_QUALITY:
b_dir = Path(self.settings.conflicts.backup_dir or ".backup") / batch_id
b_dir.mkdir(parents=True, exist_ok=True)
backup_dst = b_dir / dst.name
shutil.move(dst, backup_dst)
backup_path = str(backup_dst)
elif not self.settings.conflicts.allow_overwrite:
raise FileExistsError(f"Destination exists and allow_overwrite is False: {dst}")
# Create journal entry in IN_PROGRESS state
op = Operation(
batch_id=batch_id,
src=str(src),
dst=str(dst),
action=action.value,
category=item.category,
confidence=item.confidence,
src_hash=src_hash,
backup_path=backup_path,
details=item.details,
status=OperationStatus.IN_PROGRESS.value,
)
self.session.add(op)
self.session.commit()
dst.parent.mkdir(parents=True, exist_ok=True)
try:
if action == ActionType.MOVE:
self._safe_move(src, dst)
elif action == ActionType.COPY:
self._safe_copy(src, dst)
elif action == ActionType.LINK:
if dst.exists() or dst.is_symlink():
dst.unlink()
os.symlink(src, dst)
elif action == ActionType.HARDLINK:
if dst.exists():
dst.unlink()
os.link(src, dst)
# Apply custom permissions if specified
self._apply_permissions(dst)
op.status = OperationStatus.COMMITTED.value
op.completed_at = datetime.now(timezone.utc)
op.dst_hash = compute_file_hash(dst)
# Update or create FileRecord in database
self._update_file_record(dst, item)
self.session.commit()
except Exception as e:
op.status = OperationStatus.FAILED.value
op.error_message = str(e)
self.session.commit()
raise
def _safe_move(self, src: Path, dst: Path) -> None:
"""Atomic move on same filesystem, or safe temp-copy-atomic-rename cross-filesystem."""
try:
# Check if same filesystem by comparing st_dev
src_dev = src.stat().st_dev
dst_parent_dev = dst.parent.stat().st_dev
if src_dev == dst_parent_dev:
# Same device: atomic rename
os.replace(src, dst)
return
except Exception:
pass
# Cross-filesystem move:
# 1. Copy to temp file in destination directory
temp_dst = dst.parent / f".tmp_media_sorter_{uuid.uuid4().hex}_{dst.name}"
try:
shutil.copy2(src, temp_dst)
if self.settings.permissions.preserve_attributes:
copy_extended_attributes(src, temp_dst)
# Verify file size matches
if temp_dst.stat().st_size != src.stat().st_size:
raise IOError(f"Size mismatch during copy: {temp_dst.stat().st_size} != {src.stat().st_size}")
# Atomically replace into final destination
os.replace(temp_dst, dst)
# Unlink original source
src.unlink()
finally:
if temp_dst.exists():
try:
temp_dst.unlink()
except Exception:
pass
def _safe_copy(self, src: Path, dst: Path) -> None:
"""Safe copy using temporary file and atomic replace."""
temp_dst = dst.parent / f".tmp_media_sorter_{uuid.uuid4().hex}_{dst.name}"
try:
shutil.copy2(src, temp_dst)
if self.settings.permissions.preserve_attributes:
copy_extended_attributes(src, temp_dst)
if temp_dst.stat().st_size != src.stat().st_size:
raise IOError("Copy size mismatch")
os.replace(temp_dst, dst)
finally:
if temp_dst.exists():
try:
temp_dst.unlink()
except Exception:
pass
def _apply_permissions(self, path: Path) -> None:
"""Apply configured mode bits and ownership safely."""
perm_cfg = self.settings.permissions
if perm_cfg.file_mode and path.is_file():
try:
mode = int(perm_cfg.file_mode, 8)
os.chmod(path, mode)
except Exception as e:
logger.debug("Failed setting file mode", path=str(path), error=str(e))
if (perm_cfg.owner or perm_cfg.group) and hasattr(os, "chown"):
try:
import pwd
import grp
uid = -1
gid = -1
if perm_cfg.owner:
uid = int(perm_cfg.owner) if perm_cfg.owner.isdigit() else pwd.getpwnam(perm_cfg.owner).pw_uid
if perm_cfg.group:
gid = int(perm_cfg.group) if perm_cfg.group.isdigit() else grp.getgrnam(perm_cfg.group).gr_gid
os.chown(path, uid, gid)
except Exception as e:
logger.debug("Failed setting ownership", path=str(path), error=str(e))
def _update_file_record(self, final_path: Path, item: PlannedOperation) -> None:
"""Record the file in the database to prevent re-processing."""
try:
stat = final_path.stat()
rec = self.session.query(FileRecord).filter_by(path=str(final_path)).first()
if not rec:
rec = FileRecord(
path=str(final_path),
size=stat.st_size,
mtime=stat.st_mtime,
category=item.category,
confidence=item.confidence,
status="organized",
last_processed=datetime.now(timezone.utc),
)
self.session.add(rec)
else:
rec.size = stat.st_size
rec.mtime = stat.st_mtime
rec.category = item.category
rec.confidence = item.confidence
rec.status = "organized"
rec.last_processed = datetime.now(timezone.utc)
except Exception:
pass
def _record_operation(
self, batch_id: str, item: PlannedOperation, status: OperationStatus, is_dry_run: bool
) -> None:
op = Operation(
batch_id=batch_id,
src=str(item.src),
dst=str(item.dst),
action=item.action.value,
category=item.category,
confidence=item.confidence,
details=item.details,
status=status.value,
)
self.session.add(op)
self.session.commit()
def _record_quarantine(self, batch_id: str, item: PlannedOperation, is_dry_run: bool) -> None:
q = self.session.query(QuarantineRecord).filter_by(src=str(item.src)).first()
if not q:
q = QuarantineRecord(
src=str(item.src),
suggested_category=item.category,
confidence=item.confidence,
reason=item.quarantine_reason or "Low confidence or unclassifiable",
signals=item.details,
status=QuarantineStatus.PENDING.value,
)
self.session.add(q)
self.session.commit()
if not is_dry_run and self.settings.quarantine.move_to_quarantine_folder:
# Move to quarantine folder
q_dir = self.settings.get_destination_path("quarantine") / (item.quarantine_reason or "review")
q_dir.mkdir(parents=True, exist_ok=True)
dst_path = q_dir / item.src.name
try:
self._safe_move(item.src, dst_path)
q.resolved_path = str(dst_path)
self.session.commit()
except Exception as e:
logger.error("Failed moving to quarantine folder", src=str(item.src), error=str(e))
def rollback_batch(self, batch_id: Optional[str] = None) -> int:
"""Invert all COMMITTED operations in a batch, returning count of reverted files."""
query = self.session.query(BatchRecord)
if batch_id:
batch = query.filter_by(id=batch_id).first()
else:
# Default to latest non-rolled-back completed batch
batch = (
query.filter(
BatchRecord.status.in_(["COMPLETED", "PARTIAL_FAILURE", "PARTIAL_ROLLBACK"]),
BatchRecord.dry_run == False,
)
.order_by(BatchRecord.created_at.desc())
.first()
)
if not batch:
logger.warning("No qualifying batch found for rollback", requested_id=batch_id)
return 0
logger.info("Initiating rollback", batch_id=batch.id)
# Query committed operations in reverse execution order
ops = (
self.session.query(Operation)
.filter_by(batch_id=batch.id, status=OperationStatus.COMMITTED.value)
.order_by(Operation.id.desc())
.all()
)
reverted_count = 0
failed_count = 0
reverted_dest_dirs: Set[Path] = set()
for op in ops:
src = Path(op.src)
dst = Path(op.dst)
action = op.action
try:
if action == ActionType.MOVE.value:
if dst.exists():
src.parent.mkdir(parents=True, exist_ok=True)
self._safe_move(dst, src)
reverted_count += 1
reverted_dest_dirs.add(dst.parent)
# Restore backup if one was taken
if op.backup_path and Path(op.backup_path).exists():
self._safe_move(Path(op.backup_path), dst)
elif action == ActionType.COPY.value:
if dst.exists():
dst.unlink()
reverted_count += 1
reverted_dest_dirs.add(dst.parent)
elif action in (ActionType.LINK.value, ActionType.HARDLINK.value):
if dst.exists() or dst.is_symlink():
dst.unlink()
reverted_count += 1
reverted_dest_dirs.add(dst.parent)
op.status = OperationStatus.ROLLED_BACK.value
# Delete FileRecord for destination
rec = self.session.query(FileRecord).filter_by(path=str(dst)).first()
if rec:
self.session.delete(rec)
except Exception as e:
failed_count += 1
logger.error("Error reverting operation during rollback", op_id=op.id, error=str(e))
if failed_count == 0 and reverted_count > 0:
batch.status = "ROLLED_BACK"
elif reverted_count > 0:
batch.status = "PARTIAL_ROLLBACK"
else:
batch.status = "ROLLBACK_FAILED"
self.session.commit()
# Clean empty directories in destination tree
self._clean_empty_destination_dirs(reverted_dest_dirs)
return reverted_count
def _clean_empty_destination_dirs(self, dest_dirs: Set[Path]) -> None:
"""Prune empty parent folders in destination tree after rollback."""
dest_roots = {
self.settings.get_destination_path(cat).resolve()
for cat in ("movie", "tv", "anime", "music", "audiobook", "podcast", "photo", "home_video", "documentary", "quarantine")
}
dest_base = self.settings.get_destination_base_path().resolve()
dest_roots.add(dest_base)
candidate_dirs: Set[Path] = set()
for d in dest_dirs:
try:
curr = d.resolve()
while curr not in dest_roots and any(curr.is_relative_to(r) for r in dest_roots):
candidate_dirs.add(curr)
curr = curr.parent
except Exception:
continue
sorted_dirs = sorted(candidate_dirs, key=lambda p: len(p.parts), reverse=True)
for d in sorted_dirs:
if not d.exists() or not d.is_dir() or d in dest_roots:
continue
try:
entries = [
e for e in d.iterdir()
if e.name not in (".DS_Store", "Thumbs.db", "desktop.ini")
]
if not entries:
for junk in list(d.iterdir()):
try:
junk.unlink()
except Exception:
pass
d.rmdir()
except Exception:
pass
def rollback_all(self) -> int:
"""Roll back ALL completed, non-rolled-back batches in reverse chronological order."""
batches = (
self.session.query(BatchRecord)
.filter(
BatchRecord.status.in_(["COMPLETED", "PARTIAL_FAILURE", "PARTIAL_ROLLBACK"]),
BatchRecord.dry_run == False,
)
.order_by(BatchRecord.created_at.desc())
.all()
)
total_reverted = 0
for batch in batches:
total_reverted += self.rollback_batch(batch.id)
return total_reverted
def recover_interrupted_batches(self) -> int:
"""Clean up orphaned temp files and mark interrupted operations as FAILED."""
in_progress_ops = (
self.session.query(Operation)
.filter_by(status=OperationStatus.IN_PROGRESS.value)
.all()
)
recovered_count = 0
for op in in_progress_ops:
logger.warning("Found interrupted operation during crash recovery", op_id=op.id, src=op.src, dst=op.dst)
op.status = OperationStatus.FAILED.value
op.error_message = "Interrupted by system crash or process kill"
recovered_count += 1
if recovered_count > 0:
self.session.commit()
return recovered_count
def clean_empty_directories(self, moved_src_paths: List[Path]) -> int:
"""Remove empty parent directories and delete .txt / junk files in source folders after moving files.
Ascends from moved file parent folders up to, but never removing, the source root directories.
Also removes companion or orphaned .txt files left behind in source folders.
"""
source_roots = {p.resolve() for p in self.settings.get_source_paths()}
# Also include any parent roots if configured
candidate_dirs: Set[Path] = set()
for src in moved_src_paths:
try:
# Delete companion .txt file (e.g. Movie.txt alongside Movie.mkv)
comp = src.with_suffix(".txt")
if comp.is_file():
try:
comp.unlink()
logger.info("Deleted companion .txt file during cleanup", file=str(comp))
except Exception:
pass
parent = src.resolve().parent
while parent not in source_roots and any(parent.is_relative_to(root) for root in source_roots):
candidate_dirs.add(parent)
parent = parent.parent
except Exception:
continue
# Sort candidate directories deepest first (longest path / most parts first)
sorted_dirs = sorted(candidate_dirs, key=lambda d: len(d.parts), reverse=True)
removed_count = 0
for d in sorted_dirs:
if not d.exists() or not d.is_dir():
continue
# Ensure we never delete a configured source root
if d in source_roots:
continue
try:
# Delete any .txt files in candidate directories during cleanup
for child in list(d.iterdir()):
if child.is_file() and child.name.lower().endswith(".txt"):
try:
child.unlink()
logger.info("Deleted .txt file during cleanup", file=str(child))
except Exception:
pass
# Check if directory contains any remaining files or subdirs (ignoring OS junk and .txt files)
entries = [
e for e in d.iterdir()
if e.name not in (".DS_Store", "Thumbs.db", "desktop.ini") and not e.name.lower().endswith(".txt")
]
if not entries:
# Clean up junk files before rmdir
for junk in d.iterdir():
try:
junk.unlink()
except Exception:
pass
d.rmdir()
removed_count += 1
logger.info("Cleaned up empty source directory", directory=str(d))
except (OSError, PermissionError) as e:
logger.debug("Could not remove directory (not empty or permissions issue)", directory=str(d), error=str(e))
# Also clean up any .txt files left in source roots
for root in source_roots:
if root.exists() and root.is_dir():
try:
for child in list(root.iterdir()):
if child.is_file() and child.name.lower().endswith(".txt"):
try:
child.unlink()
logger.info("Deleted .txt file in source root during cleanup", file=str(child))
except Exception:
pass
except Exception:
pass
return removed_count
+299
View File
@@ -0,0 +1,299 @@
"""Library management and show/movie memory indexing engine.
Tracks known shows and movies in the library, syncs filesystem library directories,
and provides automatic show memory routing for incoming downloads.
"""
from __future__ import annotations
import os
import re
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
import structlog
from sqlalchemy.orm import Session
from .config import Settings
from .models import LibraryItem, utc_now
logger = structlog.get_logger(__name__)
VIDEO_EXTENSIONS = {".mkv", ".mp4", ".m4v", ".avi", ".mov", ".webm", ".ts", ".flv", ".wmv"}
def clean_show_title(raw: str) -> str:
"""Strip release tags, quality, year suffixes, and bracketed text from a directory name."""
clean = re.sub(r"\[[^\]]+\]|\([^\)]+\)", "", raw).strip()
# Strip trailing quality / encoding specs
clean = re.sub(
r"(?i)\b(1080p|720p|2160p|4k|bluray|bdrip|webrip|web-dl|x264|x265|hevc|h\.?264|h\.?265|dts|aac|ac3|remux|repack)\b.*",
"",
clean,
)
# Strip Season pack identifiers like S01-S08 or Season 1
clean = re.sub(r"(?i)\b(?:s\d+[-_s\d]*|season\s*\d+.*)\b", "", clean)
clean = re.sub(r"[\._]+", " ", clean).strip(" -_")
return clean if len(clean) >= 2 else raw.strip()
def sync_library_from_disk(session: Session, settings: Settings) -> Dict[str, int]:
"""Scan configured SHOWS_DIR and MOVIES_DIR on disk and synchronize library_items."""
shows_dir = settings.get_destination_path("tv")
movies_dir = settings.get_destination_path("movie")
shows_count = 0
movies_count = 0
# Cache existing records in memory by (title.lower(), category)
existing_items: Dict[Tuple[str, str], LibraryItem] = {
(item.title.lower(), item.category): item for item in session.query(LibraryItem).all()
}
# 1. Scan Shows Directory
if shows_dir.exists() and shows_dir.is_dir():
try:
for entry in shows_dir.iterdir():
if entry.name.startswith(".") or not entry.is_dir():
continue
folder_name = entry.name
title = clean_show_title(folder_name)
if not title:
continue
# Count video files and detect seasons
episodes = 0
seasons = set()
try:
for root, _, files in os.walk(entry):
for f in files:
ext = os.path.splitext(f)[1].lower()
if ext in VIDEO_EXTENSIONS:
episodes += 1
s_m = re.search(r"(?i)\b(?:season|s)\s*(\d{1,2})\b", Path(root).name)
if s_m:
seasons.add(int(s_m.group(1)))
except Exception:
pass
key = (title.lower(), "tv")
if key in existing_items:
item = existing_items[key]
item.destination_folder = str(entry)
item.item_count = max(item.item_count, episodes)
item.seasons_count = max(item.seasons_count, len(seasons))
item.last_updated = utc_now()
else:
item = LibraryItem(
title=title,
category="tv",
destination_folder=str(entry),
item_count=episodes,
seasons_count=len(seasons),
first_detected=utc_now(),
last_updated=utc_now(),
)
session.add(item)
existing_items[key] = item
shows_count += 1
except Exception as e:
logger.error("Error scanning shows directory for library", error=str(e))
# 2. Scan Movies Directory
if movies_dir.exists() and movies_dir.is_dir():
try:
for entry in movies_dir.iterdir():
if entry.name.startswith("."):
continue
title = entry.name
year = None
y_m = re.search(r"\b(19\d\d|20\d\d)\b", entry.name)
if y_m:
year = int(y_m.group(1))
title = entry.name[: y_m.start()].strip(" (.-_")
clean = clean_show_title(title)
item_files = 1
if entry.is_dir():
try:
item_files = sum(
1 for _, _, files in os.walk(entry)
for f in files if os.path.splitext(f)[1].lower() in VIDEO_EXTENSIONS
)
except Exception:
pass
key = (clean.lower(), "movie")
if key in existing_items:
item = existing_items[key]
item.destination_folder = str(entry)
item.year = year or item.year
item.item_count = max(item.item_count, item_files)
item.last_updated = utc_now()
else:
item = LibraryItem(
title=clean,
category="movie",
year=year,
destination_folder=str(entry),
item_count=item_files,
first_detected=utc_now(),
last_updated=utc_now(),
)
session.add(item)
existing_items[key] = item
movies_count += 1
except Exception as e:
logger.error("Error scanning movies directory for library", error=str(e))
session.commit()
logger.info("Library synchronized with disk", shows=shows_count, movies=movies_count)
return {"shows_synced": shows_count, "movies_synced": movies_count}
def record_detected_item(
session: Session,
settings: Settings,
title: str,
category: str,
destination_folder: Optional[str] = None,
year: Optional[int] = None,
poster_url: Optional[str] = None,
delta_count: int = 0,
) -> LibraryItem:
"""Record or update a show or movie in the library database."""
category = category.lower()
if category not in ("tv", "movie"):
category = "tv"
clean = clean_show_title(title) if category == "tv" else title.strip()
item = session.query(LibraryItem).filter_by(title=clean, category=category).first()
if not destination_folder:
dest_base = settings.get_destination_path(category)
destination_folder = str(dest_base / clean)
if item:
if destination_folder:
item.destination_folder = destination_folder
if year:
item.year = year
if poster_url and not item.poster_url:
item.poster_url = poster_url
if delta_count:
item.item_count = max(0, item.item_count + delta_count)
item.last_updated = utc_now()
else:
item = LibraryItem(
title=clean,
category=category,
year=year,
destination_folder=destination_folder,
poster_url=poster_url,
item_count=max(0, delta_count),
first_detected=utc_now(),
last_updated=utc_now(),
)
session.add(item)
session.commit()
return item
def get_known_shows(session: Session) -> List[Dict[str, Any]]:
"""Return all known TV shows in the library for matching."""
items = session.query(LibraryItem).filter_by(category="tv").all()
shows = []
for item in items:
clean = item.title.strip()
if len(clean) >= 2:
shows.append({
"title": clean,
"raw_title": item.title,
"destination_folder": item.destination_folder,
"poster_url": item.poster_url,
"item_count": item.item_count,
})
# Sort by title length descending so longer specific titles match first
shows.sort(key=lambda x: len(x["title"]), reverse=True)
return shows
def match_known_show(filename_or_text: str, known_shows: List[Dict[str, Any]]) -> Optional[Dict[str, Any]]:
"""Check if filename_or_text contains or matches a known show in the library."""
if not filename_or_text or not known_shows:
return None
# Replace separators with spaces
normalized = re.sub(r"[\._]+", " ", filename_or_text)
for show in known_shows:
title = show["title"]
if len(title) < 3:
continue
# Check whole word match
pattern = r"(?i)(?<![a-z0-9])" + re.escape(title) + r"(?![a-z0-9])"
if re.search(pattern, normalized):
return show
return None
def list_library_items(
session: Session,
category: Optional[str] = None,
search: Optional[str] = None,
) -> Dict[str, Any]:
"""List library items with counts, optionally filtered by category and search term."""
q = session.query(LibraryItem)
if category and category.lower() in ("tv", "movie"):
q = q.filter_by(category=category.lower())
if search:
s = f"%{search.strip()}%"
q = q.filter(LibraryItem.title.ilike(s))
items = q.order_by(LibraryItem.title.asc()).all()
total_shows = session.query(LibraryItem).filter_by(category="tv").count()
total_movies = session.query(LibraryItem).filter_by(category="movie").count()
shows_list = []
movies_list = []
for item in items:
d = {
"id": item.id,
"title": item.title,
"category": item.category,
"year": item.year,
"destination_folder": item.destination_folder,
"poster_url": item.poster_url,
"item_count": item.item_count,
"seasons_count": item.seasons_count,
"first_detected": item.first_detected.isoformat() if item.first_detected else None,
"last_updated": item.last_updated.isoformat() if item.last_updated else None,
}
if item.category == "tv":
shows_list.append(d)
else:
movies_list.append(d)
return {
"total_shows": total_shows,
"total_movies": total_movies,
"shows": shows_list,
"movies": movies_list,
}
def clear_library(session: Session) -> int:
"""Clear all indexed show and movie items from the library catalog database."""
deleted_count = session.query(LibraryItem).delete()
session.commit()
logger.info("Library catalog cleared", deleted_count=deleted_count)
return deleted_count
+184
View File
@@ -0,0 +1,184 @@
"""SQLAlchemy database models for Media Sorter.
Provides data structures for file tracking, operation journaling, quarantine,
batch execution, and configuration auditing.
"""
from __future__ import annotations
import datetime
from enum import Enum
from typing import Any, Dict, Optional
from sqlalchemy import (
Boolean,
Column,
DateTime,
Float,
ForeignKey,
Index,
Integer,
JSON,
String,
Text,
UniqueConstraint,
)
from sqlalchemy.orm import declarative_base, relationship
Base = declarative_base()
class OperationStatus(str, Enum):
PLANNED = "PLANNED"
IN_PROGRESS = "IN_PROGRESS"
COMMITTED = "COMMITTED"
FAILED = "FAILED"
ROLLED_BACK = "ROLLED_BACK"
SKIPPED = "SKIPPED"
class QuarantineStatus(str, Enum):
PENDING = "PENDING"
RESOLVED = "RESOLVED"
IGNORED = "IGNORED"
def utc_now():
return datetime.datetime.now(datetime.timezone.utc)
class BatchRecord(Base):
"""Tracks an execution batch (one invocation of media-sorter run or dry-run)."""
__tablename__ = "batches"
id = Column(String(36), primary_key=True) # UUIDv4
created_at = Column(DateTime(timezone=True), default=utc_now, nullable=False)
completed_at = Column(DateTime(timezone=True), nullable=True)
dry_run = Column(Boolean, default=False, nullable=False)
status = Column(String(32), default="IN_PROGRESS", nullable=False) # IN_PROGRESS, COMPLETED, FAILED, ROLLED_BACK
total_files = Column(Integer, default=0, nullable=False)
moved_files = Column(Integer, default=0, nullable=False)
skipped_files = Column(Integer, default=0, nullable=False)
failed_files = Column(Integer, default=0, nullable=False)
quarantined_files = Column(Integer, default=0, nullable=False)
operations = relationship("Operation", back_populates="batch", cascade="all, delete-orphan")
__table_args__ = (
Index("ix_batches_created_at", "created_at"),
)
class FileRecord(Base):
"""Tracks known files to avoid redundant probing and detect changes."""
__tablename__ = "files"
id = Column(Integer, primary_key=True, autoincrement=True)
path = Column(String(1024), unique=True, nullable=False)
size = Column(Integer, nullable=False)
mtime = Column(Float, nullable=False)
content_hash = Column(String(64), nullable=True)
status = Column(String(32), default="scanned", nullable=False) # scanned, organized, quarantined, skipped, error
category = Column(String(32), nullable=True)
confidence = Column(Float, nullable=True)
first_seen = Column(DateTime(timezone=True), default=utc_now, nullable=False)
last_processed = Column(DateTime(timezone=True), nullable=True)
__table_args__ = (
Index("ix_files_path", "path"),
Index("ix_files_status", "status"),
)
class Operation(Base):
"""Operation journal entry for atomic moves, copies, or links."""
__tablename__ = "operations"
id = Column(Integer, primary_key=True, autoincrement=True)
batch_id = Column(String(36), ForeignKey("batches.id"), nullable=False)
src = Column(String(1024), nullable=False)
dst = Column(String(1024), nullable=False)
action = Column(String(32), nullable=False) # move, copy, link, hardlink
status = Column(String(32), default=OperationStatus.PLANNED.value, nullable=False)
category = Column(String(32), nullable=True)
confidence = Column(Float, nullable=True)
src_hash = Column(String(64), nullable=True)
dst_hash = Column(String(64), nullable=True)
backup_path = Column(String(1024), nullable=True)
details = Column(JSON, nullable=True) # reasoning, sidecars, format info
error_message = Column(Text, nullable=True)
created_at = Column(DateTime(timezone=True), default=utc_now, nullable=False)
completed_at = Column(DateTime(timezone=True), nullable=True)
batch = relationship("BatchRecord", back_populates="operations")
__table_args__ = (
Index("ix_operations_batch_id", "batch_id"),
Index("ix_operations_status", "status"),
Index("ix_operations_src", "src"),
Index("ix_operations_dst", "dst"),
)
class QuarantineRecord(Base):
"""Stores files that failed confidence threshold or require manual user review."""
__tablename__ = "quarantine"
id = Column(Integer, primary_key=True, autoincrement=True)
src = Column(String(1024), unique=True, nullable=False)
suggested_category = Column(String(32), nullable=True)
confidence = Column(Float, nullable=True)
reason = Column(String(256), nullable=False)
signals = Column(JSON, nullable=True) # Diagnostic details of why it was flagged
status = Column(String(32), default=QuarantineStatus.PENDING.value, nullable=False)
resolved_path = Column(String(1024), nullable=True)
created_at = Column(DateTime(timezone=True), default=utc_now, nullable=False)
resolved_at = Column(DateTime(timezone=True), nullable=True)
__table_args__ = (
Index("ix_quarantine_status", "status"),
Index("ix_quarantine_src", "src"),
)
class ConfigAudit(Base):
"""Tracks configuration states for reproducibility and auditing."""
__tablename__ = "config_audit"
id = Column(Integer, primary_key=True, autoincrement=True)
loaded_at = Column(DateTime(timezone=True), default=utc_now, nullable=False)
config_json = Column(JSON, nullable=False)
__table_args__ = (
Index("ix_config_loaded", "loaded_at"),
)
class LibraryItem(Base):
"""Tracks known shows and movies in the user's library for automated routing and cataloging."""
__tablename__ = "library_items"
id = Column(Integer, primary_key=True, autoincrement=True)
title = Column(String(256), nullable=False)
category = Column(String(32), nullable=False) # "tv" or "movie"
year = Column(Integer, nullable=True)
destination_folder = Column(String(1024), nullable=False)
poster_url = Column(String(1024), nullable=True)
item_count = Column(Integer, default=0, nullable=False)
seasons_count = Column(Integer, default=0, nullable=False)
first_detected = Column(DateTime(timezone=True), default=utc_now, nullable=False)
last_updated = Column(DateTime(timezone=True), default=utc_now, nullable=False)
extra_info = Column(JSON, nullable=True)
__table_args__ = (
UniqueConstraint("title", "category", name="uq_library_title_category"),
Index("ix_library_category", "category"),
Index("ix_library_title", "title"),
)
+413
View File
@@ -0,0 +1,413 @@
"""Naming and path formatting engine for Media Sorter.
Renders user-defined naming templates, safely formats multi-part tags, pairs sidecars
with primary media files, and enforces rigorous cross-platform filename sanitization
(Linux, Windows, macOS, NTFS, SMB/NFS, exFAT).
"""
from __future__ import annotations
import os
import re
import unicodedata
from pathlib import Path
from typing import Any, Dict, Optional
from .classifier import ClassificationResult
from .config import Settings
# Windows reserved device names
RESERVED_NAMES = {
"CON", "PRN", "AUX", "NUL",
"COM1", "COM2", "COM3", "COM4", "COM5", "COM6", "COM7", "COM8", "COM9",
"LPT1", "LPT2", "LPT3", "LPT4", "LPT5", "LPT6", "LPT7", "LPT8", "LPT9",
}
# Illegal characters across file systems (< > : " / \ | ? *)
FORBIDDEN_CHARS_PATTERN = re.compile(r'[<>:"/\\|?*\x00-\x1f]')
def sanitize_filename_component(name: str, max_length: int = 240) -> str:
"""Sanitize an individual filename or folder name component for safe cross-platform use."""
# 1. Unicode normalization (NFC)
clean = unicodedata.normalize("NFC", name)
# 2. Replace forbidden characters with safe hyphen or space
clean = FORBIDDEN_CHARS_PATTERN.sub("-", clean)
# 3. Collapse multiple whitespace and hyphens
clean = re.sub(r"\s+", " ", clean)
clean = re.sub(r"-{2,}", "-", clean)
# 4. Strip leading/trailing spaces, dots, and hyphens (vital for Windows / SMB)
clean = clean.strip(" .-")
if not clean:
clean = "unnamed"
# 5. Check Windows reserved words
upper_base = clean.split(".")[0].upper()
if upper_base in RESERVED_NAMES:
clean = f"_{clean}"
# 6. Truncate byte length for filesystem limits (e.g. 255 bytes on ext4/NTFS/ZFS)
encoded = clean.encode("utf-8")
if len(encoded) > max_length:
parts = clean.rsplit(".", 1)
if len(parts) == 2 and 1 <= len(parts[1]) <= 10:
base, ext = parts
ext_bytes = len(f".{ext}".encode("utf-8"))
avail = max(max_length - ext_bytes, 10)
base_enc = base.encode("utf-8")[:avail]
base_clean = base_enc.decode("utf-8", errors="ignore").rstrip(" .-")
clean = f"{base_clean}.{ext}" if base_clean else ext
else:
clean = encoded[:max_length].decode("utf-8", errors="ignore").rstrip(" .-")
clean = clean.strip(" .-")
return clean if clean else "unnamed"
DEFAULT_TEMPLATES = {
"tv": "{title}/Season {season:02d}/{show_name}_{season_episode}.{ext}",
"movie": "{title} ({year})/{movie_name}.{ext}",
"anime": "{title}/Season {season:02d}/{show_name}_{season_episode} [{group}].{ext}",
}
class MediaNamer:
"""Renders organized destination paths from templates and classification results."""
def __init__(self, settings: Settings):
self.settings = settings
def generate_destination_path(
self,
cls_result: ClassificationResult,
primary_dst_path: Optional[Path] = None,
) -> Path:
"""Construct full destination path for a given file and its classification."""
category = cls_result.category
base_dir = self.settings.get_destination_path(category)
src_path = cls_result.metadata.path if cls_result.metadata else Path("unknown")
ext = src_path.suffix.lstrip(".")
# Handle Quarantine routing
if cls_result.needs_quarantine or category == "unknown":
reason = cls_result.quarantine_reason or "low_confidence"
safe_reason = sanitize_filename_component(reason)
q_template = self.settings.templates.quarantine
filename = sanitize_filename_component(src_path.name)
rel_str = q_template.format(reason=safe_reason, filename=filename, ext=ext)
return (self.settings.get_destination_path("quarantine") / rel_str).resolve()
# Handle Sidecars (Subtitles, Artwork, Metadata, Extras)
if category in ("subtitle", "artwork", "metadata"):
return self._format_sidecar_path(cls_result, primary_dst_path, base_dir)
# Retrieve template
template = getattr(self.settings.templates, category, None)
context = self._build_context(cls_result)
if category == "tv" and (not template or template == DEFAULT_TEMPLATES.get("tv")):
formatted_rel = self._format_tv_path(cls_result, context)
elif category == "movie" and (not template or template == DEFAULT_TEMPLATES.get("movie")):
formatted_rel = self._format_movie_path(cls_result, context)
elif category == "anime" and (not template or template == DEFAULT_TEMPLATES.get("anime")):
formatted_rel = self._format_anime_path(cls_result, context)
elif category == "podcast" and (not template or template == "{show}/{year}/{show} - {date} - {title}.{ext}"):
formatted_rel = self._format_podcast_path(cls_result, context)
else:
if not template:
template = "{filename}.{ext}"
formatted_rel = self._render_template(template, context)
# If file renaming is disabled, preserve original source filename
if not getattr(self.settings.general, "rename_files", True) and src_path.name != "unknown":
rel_path = Path(formatted_rel)
if len(rel_path.parts) > 1:
formatted_rel = str(rel_path.parent / src_path.name)
else:
formatted_rel = src_path.name
# Sanitize each path component separately to preserve folder hierarchy
parts = Path(formatted_rel).parts
sanitized_parts = [sanitize_filename_component(p) for p in parts]
return (base_dir / Path(*sanitized_parts)).resolve()
def _format_tv_path(self, cls_result: ClassificationResult, context: Dict[str, Any]) -> str:
show_name = context["show_name"]
ext = context["ext"]
tokens = cls_result.tokens
# Daily / dated broadcast TV formatting
date_val = context.get("date_val")
if (tokens and tokens.is_daily) or (date_val and (not tokens or not tokens.season or tokens.season > 1000)):
year = context.get("year")
if not year or year == "Unknown":
year = date_val.split("-")[0] if date_val else "Unknown"
return f"{show_name}/Season {year}/{show_name} - {date_val}.{ext}"
# Standard TV formatting (supporting Season 00, multi-ep, and season pack)
season_num = context["season"]
season_folder = f"Season {season_num:02d}"
season_episode = context["season_episode"]
return f"{show_name}/{season_folder}/{show_name} - {season_episode}.{ext}"
def _format_movie_path(self, cls_result: ClassificationResult, context: Dict[str, Any]) -> str:
title = context["title"]
year = context["year"]
ext = context["ext"]
has_year = year and year != "Unknown"
folder_name = f"{title} ({year})" if has_year else title
base_name = f"{title} ({year})" if has_year else title
edition_tag = context.get("edition_tag", "")
part_tag = context.get("part_tag", "")
extra_tag = context.get("extra_tag", "")
return f"{folder_name}/{base_name}{edition_tag}{part_tag}{extra_tag}.{ext}"
def _format_anime_path(self, cls_result: ClassificationResult, context: Dict[str, Any]) -> str:
title = context["title"]
ext = context["ext"]
tokens = cls_result.tokens
group_tag = context.get("group_tag", "")
# Multi-episode anime
if tokens and tokens.multi_episodes and len(tokens.multi_episodes) >= 2:
first_ep = tokens.multi_episodes[0]
last_ep = tokens.multi_episodes[-1]
ep_str = f"{first_ep:02d}-{last_ep:02d}"
return f"{title}/{title} - {ep_str}{group_tag}.{ext}"
# Single episode anime
if tokens and tokens.episode is not None:
ep = tokens.episode
ep_str = f"{ep:02d}" if ep < 10 else str(ep)
return f"{title}/{title} - {ep_str}{group_tag}.{ext}"
# Anime movie or special without episode number
return f"{title}/{title}{group_tag}.{ext}"
def _format_podcast_path(self, cls_result: ClassificationResult, context: Dict[str, Any]) -> str:
show = context.get("show") or context.get("artist") or "Unknown Show"
year = context.get("year")
date = context.get("date")
title = context.get("title")
ext = context.get("ext")
if title and title != show and title != "Unknown":
return f"{show}/{year}/{show} - {date} - {title}.{ext}"
return f"{show}/{year}/{show} - {date}.{ext}"
def _format_sidecar_path(
self,
cls_result: ClassificationResult,
primary_dst_path: Optional[Path],
base_dir: Path,
) -> Path:
src_path = cls_result.metadata.path
ext = src_path.suffix.lstrip(".")
if primary_dst_path:
parent_dir = primary_dst_path.parent
primary_stem = primary_dst_path.stem
if cls_result.category == "subtitle":
# Detect language code or compound tag in subtitle (e.g. movie.en.srt, movie.forced.srt)
src_stem = src_path.stem
m = re.search(
r"\.((?:[a-zA-Z]{2,3}\.)?(?:forced|sdh|cc)|[a-zA-Z]{2,3}(?:-[a-zA-Z]{2,4})?)$",
src_stem,
re.IGNORECASE,
)
if m:
lang_suffix = f".{m.group(1)}"
else:
parts = src_stem.split(".")
if len(parts) > 1 and len(parts[-1]) in (2, 3, 6):
lang_suffix = f".{parts[-1]}"
else:
lang_suffix = ""
new_filename = f"{primary_stem}{lang_suffix}.{ext}"
return parent_dir / sanitize_filename_component(new_filename)
elif cls_result.category == "artwork":
# e.g. poster.jpg, cover.jpg in the same movie/show folder
return parent_dir / sanitize_filename_component(src_path.name)
elif cls_result.category == "metadata":
# NFO file matches primary stem or stays alongside
new_filename = f"{primary_stem}.{ext}"
return parent_dir / sanitize_filename_component(new_filename)
# If orphan sidecar (no primary matched), place into respective folder
sanitized_name = sanitize_filename_component(src_path.name)
return (base_dir / sanitized_name).resolve()
def _build_context(self, res: ClassificationResult) -> Dict[str, Any]:
tokens = res.tokens
meta = res.metadata
src_path = meta.path if meta else Path("file")
# Fix Season 00 / Episode 00 falsy bug
season_num = tokens.season if (tokens and tokens.season is not None) else 1
episode_num = tokens.episode if (tokens and tokens.episode is not None) else 1
# Format season_episode string with multi-episode and season pack support
if tokens and tokens.multi_episodes and len(tokens.multi_episodes) >= 2:
season_ep_str = f"S{season_num:02d}E{tokens.multi_episodes[0]:02d}-E{tokens.multi_episodes[-1]:02d}"
elif tokens and (tokens.is_season_pack or (tokens.season is not None and tokens.episode is None and not getattr(tokens, "multi_episodes", None))):
season_ep_str = f"Season {season_num:02d}"
else:
season_ep_str = f"S{season_num:02d}E{episode_num:02d}"
main_title = (tokens.title if tokens else None) or src_path.stem
# Clean release group: omit when unknown, NEVER emit 'UnknownGroup'
group_val = tokens.group if (tokens and tokens.group and tokens.group != "UnknownGroup") else ""
group_tag = f" [{group_val}]" if group_val else ""
# Extract movie edition, part, and extra tags
edition_val = getattr(tokens, "edition", None) if tokens else None
if not edition_val:
em = re.search(r"\b(extended|directors?\.cut|remastered|criterion(?:\.collection)?|final\.cut)\b", src_path.stem, re.I)
if em:
raw_ed = em.group(1).lower().replace(".", " ")
if "director" in raw_ed:
edition_val = "Director's Cut"
elif "criterion" in raw_ed:
edition_val = "Criterion"
elif "final" in raw_ed:
edition_val = "Final Cut"
elif "remaster" in raw_ed:
edition_val = "Remastered"
elif "extend" in raw_ed:
edition_val = "Extended"
edition_tag = f" [{edition_val}]" if edition_val else ""
part_val = getattr(tokens, "part", None) if tokens else None
part_label = getattr(tokens, "part_label", None) if tokens else None
if part_val is None:
pm = re.search(r"\b(?:cd|part|pt)[\.\s_-]*(\d+)\b", src_path.stem, re.I)
if pm:
part_val = int(pm.group(1))
part_label = f"Pt.{part_val}"
elif not part_label:
part_label = f"Pt.{part_val}"
part_tag = f" [{part_label}]" if part_label else ""
extra_m = re.search(r"-(behindthescenes|deleted|trailer|featurette)\b", src_path.stem, re.I)
extra_tag = f"-{extra_m.group(1).lower()}" if extra_m else ""
# Date resolution for daily TV shows and podcasts
date_val = getattr(tokens, "air_date", None) or (tokens.date_stamp if tokens and not tokens.is_photo_or_home_video else None)
if not date_val:
dm = re.search(r"\b((?:19|20)\d{2})[-._](0[1-9]|1[0-2])[-._](0[1-9]|[12]\d|3[01])\b", src_path.stem)
if dm:
date_val = f"{dm.group(1)}-{dm.group(2)}-{dm.group(3)}"
ctx: Dict[str, Any] = {
"ext": src_path.suffix.lstrip("."),
"filename": src_path.stem,
"title": main_title,
"show_name": main_title,
"SHOW_NAME": main_title,
"movie_name": main_title,
"MOVIE_NAME": main_title,
"season_episode": season_ep_str,
"SEASON_EPISODE": season_ep_str,
"year": (tokens.year if tokens else None) or "Unknown",
"season": season_num,
"episode": episode_num,
"episode_title": (tokens.episode_title if tokens else None) or f"Episode {episode_num}",
"artist": (tokens.artist if tokens else None) or "Unknown Artist",
"album": (tokens.album if tokens else None) or "Unknown Album",
"track": (tokens.track if tokens else 1) or 1,
"disc": (tokens.disc if tokens else 1) or 1,
"group": group_val,
"group_tag": group_tag,
"edition": edition_val,
"edition_tag": edition_tag,
"part": part_val,
"part_label": part_label,
"part_tag": part_tag,
"extra_tag": extra_tag,
"date_val": date_val,
"resolution": (tokens.resolution if tokens and tokens.resolution else (meta.resolution_label if meta else "")),
"codec": (tokens.video_codec or (meta.codec_video if meta else "h264")),
"author": (tokens.artist if tokens else None) or "Unknown Author",
"chapter": (tokens.title if tokens else None) or f"Chapter {tokens.track if tokens else 1}",
"show": (tokens.artist if tokens else None) or "Unknown Show",
"date": date_val or ((tokens.date_stamp if tokens else "2026-01-01") or "2026-01-01"),
"month": 1,
"day": 1,
"time": "000000",
"camera": "Camera",
"event": "Event",
}
# Override from provider result if available
if res.provider_result:
p = res.provider_result
if p.canonical_title:
ctx["title"] = p.canonical_title
ctx["show_name"] = p.canonical_title
ctx["SHOW_NAME"] = p.canonical_title
ctx["movie_name"] = p.canonical_title
ctx["MOVIE_NAME"] = p.canonical_title
if p.year:
ctx["year"] = p.year
if p.episode_title:
ctx["episode_title"] = p.episode_title
if p.artist:
ctx["artist"] = p.artist
if p.album:
ctx["album"] = p.album
# Parse date stamp fields if present
date_source = (tokens and tokens.date_stamp) or (ctx.get("date") if res.category in ("home_video", "photo", "podcast") else None)
if date_source and date_source != "Unknown":
date_parts = str(date_source).split("-")
if len(date_parts) == 3:
try:
if ctx["year"] == "Unknown":
ctx["year"] = int(date_parts[0])
ctx["month"] = int(date_parts[1])
ctx["day"] = int(date_parts[2])
except ValueError:
pass
# Clean tags from meta
if meta and meta.tags:
if "camera_model" in meta.tags:
ctx["camera"] = sanitize_filename_component(meta.tags["camera_model"])
if "album" in meta.tags and not ctx.get("album"):
ctx["album"] = meta.tags["album"]
if "artist" in meta.tags and not ctx.get("artist"):
ctx["artist"] = meta.tags["artist"]
return ctx
def _render_template(self, template: str, context: Dict[str, Any]) -> str:
"""Format template while gracefully cleaning empty technical brackets."""
# Normalize <TAG> to {TAG} for convenience if users use angle brackets
rendered = re.sub(r"<([a-zA-Z_0-9]+)>", r"{\1}", template)
try:
rendered = rendered.format(**context)
except (KeyError, ValueError):
# Safe token replacement if format specifier fails
safe_ctx = {k: str(v) if v is not None else "" for k, v in context.items()}
# Remove format specifiers like :02d
simplified = re.sub(r"\{(\w+):[^}]+\}", r"{\1}", rendered)
try:
rendered = simplified.format(**safe_ctx)
except Exception:
rendered = f"{context.get('title', 'media')}.{context.get('ext', 'bin')}"
# Clean empty technical brackets such as "[]" or "[ ]" or "()"
rendered = re.sub(r"\[\s*\]", "", rendered)
rendered = re.sub(r"\(\s*\)", "", rendered)
rendered = re.sub(r"\s{2,}", " ", rendered)
return rendered.strip()
+67
View File
@@ -0,0 +1,67 @@
"""Notification dispatcher for Media Sorter.
Sends batch summary alerts and error notices to webhook endpoints (Slack, Discord, generic JSON).
"""
from __future__ import annotations
from typing import Any, Dict, Optional
import requests
import structlog
from .config import NotificationSettings
from .executor import BatchExecutionReport
logger = structlog.get_logger(__name__)
def send_batch_notification(
settings: NotificationSettings, report: BatchExecutionReport
) -> bool:
"""Send webhook alert for completed or failed batch."""
if not settings.enabled or not settings.webhook_url:
return False
is_failure = report.failed_files > 0
if is_failure and not settings.notify_on_failure:
return False
if not is_failure and not settings.notify_on_complete:
return False
status_str = "FAILED" if is_failure else ("DRY RUN PREVIEW" if report.dry_run else "SUCCESS")
title = f"Media Sorter: {status_str} [Batch {report.batch_id[:8]}]"
summary_text = (
f"**{title}**\n"
f"• Total Files: {report.total_files}\n"
f"• Organized/Moved: {report.moved_files}\n"
f"• Copied: {report.copied_files}\n"
f"• Skipped: {report.skipped_files}\n"
f"• Quarantined: {report.quarantined_files}\n"
f"• Failures: {report.failed_files}\n"
)
if report.errors:
summary_text += f"\nErrors:\n" + "\n".join(f"- {e}" for e in report.errors[:5])
payload: Dict[str, Any] = {
"text": summary_text,
"content": summary_text, # Discord format compatibility
"batch_id": report.batch_id,
"dry_run": report.dry_run,
"status": status_str,
"total_files": report.total_files,
"moved_files": report.moved_files,
"quarantined_files": report.quarantined_files,
"failed_files": report.failed_files,
}
try:
resp = requests.post(settings.webhook_url, json=payload, timeout=5.0)
if resp.status_code in (200, 201, 204):
return True
logger.warning("Webhook dispatch failed", status=resp.status_code)
except Exception as e:
logger.warning("Failed sending notification webhook", error=str(e))
return False
+255
View File
@@ -0,0 +1,255 @@
"""Metadata provider integrations for Media Sorter.
Provides interfaces and implementations for querying external metadata (TMDB, TVDB,
MusicBrainz) with rate-limiting, request caching, and offline fallbacks.
"""
from __future__ import annotations
import hashlib
import json
import time
from abc import ABC, abstractmethod
from dataclasses import dataclass
from typing import Any, Dict, List, Optional, Tuple
import requests
import structlog
logger = structlog.get_logger(__name__)
@dataclass
class ProviderResult:
canonical_title: str
year: Optional[int] = None
media_type: str = "movie" # movie, tv, anime, music
season: Optional[int] = None
episode: Optional[int] = None
episode_title: Optional[str] = None
artist: Optional[str] = None
album: Optional[str] = None
genres: List[str] = None
confidence_boost: float = 0.15
raw_payload: Optional[Dict[str, Any]] = None
class MetadataProvider(ABC):
"""Abstract base class for all metadata providers."""
@abstractmethod
def search_movie(self, title: str, year: Optional[int] = None) -> Optional[ProviderResult]:
pass
@abstractmethod
def search_tv(self, title: str, year: Optional[int] = None, season: Optional[int] = None, episode: Optional[int] = None) -> Optional[ProviderResult]:
pass
@abstractmethod
def search_music(self, artist: str, album: Optional[str] = None, title: Optional[str] = None) -> Optional[ProviderResult]:
pass
class MemoryCache:
"""In-memory cache with TTL for metadata queries."""
def __init__(self, ttl_seconds: int = 86400):
self.ttl = ttl_seconds
self._store: Dict[str, Tuple[float, Any]] = {}
def get(self, key: str) -> Optional[Any]:
if key in self._store:
timestamp, data = self._store[key]
if time.time() - timestamp < self.ttl:
return data
del self._store[key]
return None
def set(self, key: str, data: Any) -> None:
self._store[key] = (time.time(), data)
class TMDBProvider(MetadataProvider):
"""TheMovieDatabase (TMDB) API provider with rate-limiting and caching."""
BASE_URL = "https://api.themoviedb.org/3"
def __init__(self, api_key: Optional[str] = None, rate_limit_per_second: float = 2.0, cache_ttl_seconds: int = 86400):
self.api_key = api_key
self.min_interval = 1.0 / max(rate_limit_per_second, 0.1)
self.last_request_time = 0.0
self.cache = MemoryCache(ttl_seconds=cache_ttl_seconds)
def _throttle(self) -> None:
elapsed = time.time() - self.last_request_time
if elapsed < self.min_interval:
time.sleep(self.min_interval - elapsed)
self.last_request_time = time.time()
def _query(self, endpoint: str, params: Dict[str, Any]) -> Optional[Dict[str, Any]]:
if not self.api_key:
return None
cache_key = f"tmdb:{endpoint}:{json.dumps(params, sort_keys=True)}"
cached = self.cache.get(cache_key)
if cached is not None:
return cached
self._throttle()
req_params = dict(params)
req_params["api_key"] = self.api_key
try:
resp = requests.get(f"{self.BASE_URL}/{endpoint}", params=req_params, timeout=5.0)
if resp.status_code == 200:
data = resp.json()
self.cache.set(cache_key, data)
return data
logger.warning("TMDB request failed", status=resp.status_code, endpoint=endpoint)
except Exception as e:
logger.warning("TMDB network error", error=str(e))
return None
def search_movie(self, title: str, year: Optional[int] = None) -> Optional[ProviderResult]:
params: Dict[str, Any] = {"query": title}
if year:
params["year"] = year
data = self._query("search/movie", params)
if not data or not data.get("results"):
return None
first = data["results"][0]
release_date = first.get("release_date", "")
res_year = int(release_date[:4]) if len(release_date) >= 4 and release_date[:4].isdigit() else year
return ProviderResult(
canonical_title=first.get("title", title),
year=res_year,
media_type="movie",
confidence_boost=0.15,
raw_payload=first,
)
def search_tv(self, title: str, year: Optional[int] = None, season: Optional[int] = None, episode: Optional[int] = None) -> Optional[ProviderResult]:
params: Dict[str, Any] = {"query": title}
if year:
params["first_air_date_year"] = year
data = self._query("search/tv", params)
if not data or not data.get("results"):
return None
first = data["results"][0]
show_id = first.get("id")
show_title = first.get("name", title)
air_date = first.get("first_air_date", "")
res_year = int(air_date[:4]) if len(air_date) >= 4 and air_date[:4].isdigit() else year
ep_title = None
if show_id and season is not None and episode is not None:
ep_data = self._query(f"tv/{show_id}/season/{season}/episode/{episode}", {})
if ep_data:
ep_title = ep_data.get("name")
return ProviderResult(
canonical_title=show_title,
year=res_year,
media_type="tv",
season=season,
episode=episode,
episode_title=ep_title,
confidence_boost=0.20,
raw_payload=first,
)
def search_music(self, artist: str, album: Optional[str] = None, title: Optional[str] = None) -> Optional[ProviderResult]:
return None # TMDB does not index music
class MusicBrainzProvider(MetadataProvider):
"""MusicBrainz WS2 API provider with courteous rate-limiting (1 req/sec)."""
BASE_URL = "https://musicbrainz.org/ws/2"
def __init__(self, rate_limit_per_second: float = 1.0, cache_ttl_seconds: int = 86400):
self.min_interval = 1.0 / max(rate_limit_per_second, 0.1)
self.last_request_time = 0.0
self.cache = MemoryCache(ttl_seconds=cache_ttl_seconds)
def _throttle(self) -> None:
elapsed = time.time() - self.last_request_time
if elapsed < self.min_interval:
time.sleep(self.min_interval - elapsed)
self.last_request_time = time.time()
def search_movie(self, title: str, year: Optional[int] = None) -> Optional[ProviderResult]:
return None
def search_tv(self, title: str, year: Optional[int] = None, season: Optional[int] = None, episode: Optional[int] = None) -> Optional[ProviderResult]:
return None
def search_music(self, artist: str, album: Optional[str] = None, title: Optional[str] = None) -> Optional[ProviderResult]:
query_parts = [f'artist:"{artist}"']
if album:
query_parts.append(f'release:"{album}"')
if title:
query_parts.append(f'recording:"{title}"')
query_str = " AND ".join(query_parts)
cache_key = f"mb:{query_str}"
cached = self.cache.get(cache_key)
if cached is not None:
return cached
self._throttle()
headers = {"User-Agent": "MediaSorter/0.1.0 (https://github.com/example/media-sorter)"}
params = {"query": query_str, "fmt": "json", "limit": 1}
try:
resp = requests.get(f"{self.BASE_URL}/recording", params=params, headers=headers, timeout=5.0)
if resp.status_code == 200:
data = resp.json()
recordings = data.get("recordings", [])
if recordings:
rec = recordings[0]
rec_title = rec.get("title", title or "")
# Extract release info
releases = rec.get("releases", [])
rec_album = releases[0].get("title", album) if releases else album
release_date = releases[0].get("date", "") if releases else ""
res_year = int(release_date[:4]) if len(release_date) >= 4 and release_date[:4].isdigit() else None
res = ProviderResult(
canonical_title=rec_title,
artist=artist,
album=rec_album,
year=res_year,
media_type="music",
confidence_boost=0.15,
raw_payload=rec,
)
self.cache.set(cache_key, res)
return res
except Exception as e:
logger.warning("MusicBrainz network error", error=str(e))
return None
class MockMetadataProvider(MetadataProvider):
"""Deterministic mock provider for offline testing and fixture validation."""
def __init__(self, mock_data: Optional[Dict[str, ProviderResult]] = None):
self.mock_data = mock_data or {}
def search_movie(self, title: str, year: Optional[int] = None) -> Optional[ProviderResult]:
key = f"movie:{title.lower()}"
return self.mock_data.get(key)
def search_tv(self, title: str, year: Optional[int] = None, season: Optional[int] = None, episode: Optional[int] = None) -> Optional[ProviderResult]:
key = f"tv:{title.lower()}"
return self.mock_data.get(key)
def search_music(self, artist: str, album: Optional[str] = None, title: Optional[str] = None) -> Optional[ProviderResult]:
key = f"music:{artist.lower()}"
return self.mock_data.get(key)
+141
View File
@@ -0,0 +1,141 @@
"""Quarantine and manual review management for Media Sorter.
Provides querying, manual override, reprocessing, and resolution tracking for files
that could not be safely or confidently organized automatically.
"""
from __future__ import annotations
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional
import structlog
from sqlalchemy.orm import Session
from .models import QuarantineRecord, QuarantineStatus
logger = structlog.get_logger(__name__)
class QuarantineManager:
"""Manages files held in quarantine or review status."""
def __init__(self, session: Session):
self.session = session
def list_pending(self) -> List[QuarantineRecord]:
"""Return all quarantine items waiting for human inspection."""
return (
self.session.query(QuarantineRecord)
.filter_by(status=QuarantineStatus.PENDING.value)
.order_by(QuarantineRecord.created_at.desc())
.all()
)
def get_by_id(self, item_id: int) -> Optional[QuarantineRecord]:
"""Fetch quarantine record by ID."""
return self.session.query(QuarantineRecord).filter_by(id=item_id).first()
def resolve_item(
self,
item_id: int,
resolved_category: str,
target_path: Optional[Path | str] = None,
) -> bool:
"""Resolve a quarantined item by specifying human-approved category and target path."""
rec = self.get_by_id(item_id)
if not rec:
return False
rec.status = QuarantineStatus.RESOLVED.value
rec.suggested_category = resolved_category
rec.resolved_at = datetime.now(timezone.utc)
if target_path:
rec.resolved_path = str(target_path)
self.session.commit()
logger.info("Quarantine item resolved", item_id=item_id, category=resolved_category)
return True
def ignore_item(self, item_id: int) -> bool:
"""Mark quarantine item as ignored."""
rec = self.get_by_id(item_id)
if not rec:
return False
rec.status = QuarantineStatus.IGNORED.value
rec.resolved_at = datetime.now(timezone.utc)
self.session.commit()
return True
def undo_item(self, item_id: int) -> bool:
"""Undo the resolution or quarantine status of an item.
If resolved and file was moved, moves the file back to its original src location
and resets status to PENDING. If already pending, unflags/removes from quarantine.
"""
import os
import shutil
rec = self.get_by_id(item_id)
if not rec:
return False
if rec.status == QuarantineStatus.RESOLVED.value and rec.resolved_path:
dst_path = Path(rec.resolved_path)
src_path = Path(rec.src)
if dst_path.exists():
src_path.parent.mkdir(parents=True, exist_ok=True)
try:
shutil.move(dst_path, src_path)
logger.info("Restored resolved quarantine file back to src", src=str(src_path), dst=str(dst_path))
except Exception as e:
logger.error("Failed restoring quarantine file to src", src=str(src_path), dst=str(dst_path), error=str(e))
rec.status = QuarantineStatus.PENDING.value
rec.resolved_path = None
rec.resolved_at = None
self.session.commit()
return True
elif rec.status == QuarantineStatus.PENDING.value:
# Unflag pending quarantine item
self.session.delete(rec)
self.session.commit()
return True
return False
def list_resolved(self, limit: int = 50) -> List[QuarantineRecord]:
"""Return recently resolved quarantine items."""
return (
self.session.query(QuarantineRecord)
.filter_by(status=QuarantineStatus.RESOLVED.value)
.order_by(QuarantineRecord.resolved_at.desc())
.limit(limit)
.all()
)
def get_statistics(self) -> Dict[str, int]:
"""Summarize quarantine records by status."""
total = self.session.query(QuarantineRecord).count()
pending = (
self.session.query(QuarantineRecord)
.filter_by(status=QuarantineStatus.PENDING.value)
.count()
)
resolved = (
self.session.query(QuarantineRecord)
.filter_by(status=QuarantineStatus.RESOLVED.value)
.count()
)
ignored = (
self.session.query(QuarantineRecord)
.filter_by(status=QuarantineStatus.IGNORED.value)
.count()
)
return {
"total": total,
"pending": pending,
"resolved": resolved,
"ignored": ignored,
}
+258
View File
@@ -0,0 +1,258 @@
"""Filesystem discovery and scanning engine for Media Sorter.
Efficiently traverses directories, enforces minimum file age checks (to prevent
processing files currently being written/downloaded), tests file locks, and detects
companion/sidecar files.
"""
from __future__ import annotations
import fnmatch
import os
import sys
import time
from dataclasses import dataclass, field
from pathlib import Path
from typing import Dict, Generator, List, Optional, Set, Tuple
import structlog
logger = structlog.get_logger(__name__)
# Known sidecar and companion extensions
SUBTITLE_EXTS = {".srt", ".ass", ".ssa", ".vtt", ".sub", ".idx"}
ARTWORK_NAMES = {"poster", "cover", "folder", "fanart", "banner", "clearart", "disc", "logo"}
ARTWORK_EXTS = {".jpg", ".jpeg", ".png", ".webp", ".tbn"}
METADATA_EXTS = {".nfo", ".xml", ".json"}
EXTRA_TAGS = {"-trailer", "-sample", "-featurette", "-behindthescenes", "-deleted", "-short"}
@dataclass
class ScannedFile:
path: Path
size: int
mtime: float
is_sidecar: bool = False
sidecar_type: Optional[str] = None # "subtitle", "artwork", "metadata", "extra"
primary_media_path: Optional[Path] = None
tags: Dict[str, str] = field(default_factory=dict)
def is_file_locked(path: Path) -> bool:
"""Test whether a file is currently open/locked for writing by another process."""
if not path.is_file():
return False
try:
# On Windows, try opening with exclusive read/write sharing if possible
if sys.platform == "win32":
import msvcrt
handle = open(path, "rb")
try:
msvcrt.locking(handle.fileno(), msvcrt.LK_NBLCK, 1)
msvcrt.locking(handle.fileno(), msvcrt.LK_UNLCK, 1)
finally:
handle.close()
else:
import fcntl
with open(path, "rb") as f:
fcntl.flock(f.fileno(), fcntl.LOCK_EX | fcntl.LOCK_NB)
fcntl.flock(f.fileno(), fcntl.LOCK_UN)
return False
except (IOError, OSError, PermissionError):
return True
class Scanner:
def __init__(
self,
min_file_age_seconds: int = 300,
include_patterns: Optional[List[str]] = None,
exclude_patterns: Optional[List[str]] = None,
):
self.min_file_age_seconds = min_file_age_seconds
self.include_patterns = include_patterns or ["*"]
self.exclude_patterns = exclude_patterns or [
".*",
"*.part",
"*.crdownload",
"*.!qB",
"Thumbs.db",
"desktop.ini",
"@eaDir",
"$RECYCLE.BIN",
"*.txt",
]
def _matches_filter(self, filename: str) -> bool:
"""Check whether filename matches includes and does not match excludes."""
for pattern in self.exclude_patterns:
if fnmatch.fnmatch(filename, pattern):
return False
if fnmatch.fnmatch(filename.lower(), pattern.lower()):
return False
if not self.include_patterns or "*" in self.include_patterns:
return True
for pattern in self.include_patterns:
if fnmatch.fnmatch(filename, pattern) or fnmatch.fnmatch(filename.lower(), pattern.lower()):
return True
return False
def scan_directory(self, root_dir: Path | str) -> List[ScannedFile]:
"""Recursively scan a directory returning all qualifying files."""
root = Path(root_dir).resolve()
if not root.exists():
logger.warning("Source directory does not exist", directory=str(root))
return []
now = time.time()
discovered: List[ScannedFile] = []
video_candidates: List[ScannedFile] = []
potential_sidecars: List[ScannedFile] = []
for entry_path in self._walk_safe(root):
try:
stat = entry_path.stat()
except (OSError, PermissionError) as e:
logger.warning("Skipping inaccessible file", path=str(entry_path), error=str(e))
continue
# Minimum file age check: ignore recently modified files (e.g. active downloads)
file_age = now - stat.st_mtime
if file_age < self.min_file_age_seconds:
logger.debug(
"Skipping file: modified too recently",
path=str(entry_path),
age_seconds=int(file_age),
min_age_seconds=self.min_file_age_seconds,
)
continue
# Check if locked
if is_file_locked(entry_path):
logger.debug("Skipping file: currently locked by another process", path=str(entry_path))
continue
ext = entry_path.suffix.lower()
stem = entry_path.stem.lower()
scanned = ScannedFile(
path=entry_path,
size=stat.st_size,
mtime=stat.st_mtime,
)
# Classify sidecar vs primary candidate
if ext in SUBTITLE_EXTS:
scanned.is_sidecar = True
scanned.sidecar_type = "subtitle"
potential_sidecars.append(scanned)
elif ext in ARTWORK_EXTS and any(stem == art or stem.startswith(f"{art}.") for art in ARTWORK_NAMES):
scanned.is_sidecar = True
scanned.sidecar_type = "artwork"
potential_sidecars.append(scanned)
elif ext in METADATA_EXTS:
scanned.is_sidecar = True
scanned.sidecar_type = "metadata"
potential_sidecars.append(scanned)
elif any(stem.endswith(tag) for tag in EXTRA_TAGS):
scanned.is_sidecar = True
scanned.sidecar_type = "extra"
potential_sidecars.append(scanned)
else:
discovered.append(scanned)
if ext in {".mp4", ".mkv", ".m4v", ".avi", ".mov", ".ts", ".webm"}:
video_candidates.append(scanned)
# Pair sidecars with primary files in the same directory
self._pair_sidecars(potential_sidecars, video_candidates, discovered)
return discovered
def _walk_safe(self, root: Path) -> Generator[Path, None, None]:
"""Safely traverse directories using os.scandir with cycle and permission handling."""
visited_inodes: Set[Tuple[int, int]] = set()
stack = [root]
while stack:
curr = stack.pop()
try:
with os.scandir(curr) as it:
for entry in it:
try:
# Avoid symlink loops
if entry.is_symlink():
continue
if entry.is_dir():
if not self._matches_filter(entry.name):
continue
stat = entry.stat()
dev_ino = (stat.st_dev, stat.st_ino)
if dev_ino in visited_inodes:
continue
visited_inodes.add(dev_ino)
stack.append(Path(entry.path))
elif entry.is_file():
if self._matches_filter(entry.name):
yield Path(entry.path)
except (OSError, PermissionError) as e:
logger.debug("Failed reading entry", path=entry.path, error=str(e))
except (OSError, PermissionError) as e:
logger.warning("Failed traversing directory", path=str(curr), error=str(e))
def _pair_sidecars(
self,
sidecars: List[ScannedFile],
primaries: List[ScannedFile],
all_discovered: List[ScannedFile],
) -> None:
"""Associate sidecar files (subtitles, artwork, nfo) with primary media files."""
# Index primaries by parent dir
primaries_by_dir: Dict[Path, List[ScannedFile]] = {}
for p in primaries:
primaries_by_dir.setdefault(p.path.parent.resolve(), []).append(p)
for s in sidecars:
parent = s.path.parent.resolve()
candidates = primaries_by_dir.get(parent, [])
# Support subdirectories like Subs/ or Subtitles/
if not candidates and parent.name.lower() in ("subs", "subtitles", "sub"):
parent = parent.parent
candidates = primaries_by_dir.get(parent, [])
matched_primary = None
s_stem = s.path.stem.lower()
# Sort candidates longest stem first so "Movie.Part2" matches before "Movie"
sorted_candidates = sorted(candidates, key=lambda c: len(c.path.stem), reverse=True)
for c in sorted_candidates:
c_stem = c.path.stem.lower()
if s_stem == c_stem:
matched_primary = c
break
# Check delimiter boundary: must be followed by '.', '-', '_', or ' '
if s_stem.startswith(c_stem) and len(s_stem) > len(c_stem):
next_char = s_stem[len(c_stem)]
if next_char in (".", "-", "_", " "):
matched_primary = c
break
# Fallback for single-video directories with generic sidecars (movie.nfo, poster.jpg, en.srt)
if not matched_primary and len(candidates) == 1:
if (
s.sidecar_type in ("metadata", "artwork")
or s.path.suffix.lower() in METADATA_EXTS
or s.path.suffix.lower() in SUBTITLE_EXTS
or s.path.suffix.lower() in ARTWORK_EXTS
):
matched_primary = candidates[0]
if matched_primary:
s.primary_media_path = matched_primary.path
all_discovered.append(s)
File diff suppressed because it is too large. Load diff
+323
View File
@@ -0,0 +1,323 @@
"""Main orchestration pipeline for Media Sorter.
Coordinates filesystem scanning, caching, parallel metadata probing, classification,
naming, dry-run previews, atomic execution, and audit logging.
"""
from __future__ import annotations
import concurrent.futures
import time
from pathlib import Path
from typing import Callable, Dict, List, Optional, Tuple
import structlog
from sqlalchemy.engine import Engine
from sqlalchemy.orm import Session
from .analyzer import MediaAnalyzer, MediaMetadata
from .classifier import ClassificationResult, MediaClassifier
from .config import ActionType, Settings
from .db import get_db_session
from .executor import BatchExecutionReport, MediaExecutor, PlannedOperation, acquire_process_lock
from .models import FileRecord, OperationStatus
from .namer import MediaNamer
from .notifications import send_batch_notification
from .providers import MetadataProvider, TMDBProvider
from .scanner import ScannedFile, Scanner
from .tokenizer import FilenameTokenizer, TokenizedFilename
logger = structlog.get_logger(__name__)
class MediaSorterApp:
"""High-performance orchestrator for analyzing and organizing media collections."""
def __init__(self, settings: Settings, engine: Engine):
self.settings = settings
self.engine = engine
# Component instances
self.scanner = Scanner(
min_file_age_seconds=settings.general.min_file_age_seconds,
include_patterns=settings.filters.include_patterns,
exclude_patterns=settings.filters.exclude_patterns,
)
self.tokenizer = FilenameTokenizer()
self.analyzer = MediaAnalyzer()
# Metadata provider
provider: Optional[MetadataProvider] = None
if settings.providers.enable_online_metadata:
provider = TMDBProvider(
api_key=settings.providers.tmdb_api_key,
rate_limit_per_second=settings.providers.rate_limit_per_second,
)
self.classifier = MediaClassifier(
confidence_threshold=settings.general.confidence_threshold,
provider=provider,
)
self.namer = MediaNamer(settings)
def scan_and_analyze(
self,
progress_callback: Optional[Callable[[int, int, str], None]] = None,
filter_paths: Optional[List[Path]] = None,
) -> List[Tuple[ScannedFile, ClassificationResult]]:
"""Discover files across all configured source directories and analyze in parallel."""
source_paths = self.settings.get_source_paths()
all_scanned: List[ScannedFile] = []
logger.info("Starting library discovery", sources=[str(p) for p in source_paths])
for src in source_paths:
if src.exists():
discovered = self.scanner.scan_directory(src)
all_scanned.extend(discovered)
else:
logger.warning("Configured source path does not exist", path=str(src))
logger.info("Filesystem discovery complete", total_discovered=len(all_scanned))
if not all_scanned:
return []
# Check DB cache to skip unchanged already organized files
qualifying_files: List[ScannedFile] = []
with get_db_session(self.engine) as session:
for s in all_scanned:
rec = (
session.query(FileRecord)
.filter_by(path=str(s.path), size=s.size, mtime=s.mtime, status="organized")
.first()
)
if rec:
logger.debug("Skipping unchanged already organized file", path=str(s.path))
continue
qualifying_files.append(s)
if filter_paths:
filter_resolved = {p.resolve() for p in filter_paths}
qualifying_files = [s for s in qualifying_files if s.path.resolve() in filter_resolved]
total_files = len(qualifying_files)
logger.info("Files requiring processing", count=total_files)
# Check known shows from library
known_shows: List[Dict[str, Any]] = []
try:
with get_db_session(self.engine) as session:
from .library import get_known_shows
known_shows = get_known_shows(session)
except Exception:
pass
# Parallel analysis and classification
results: List[Tuple[ScannedFile, ClassificationResult]] = []
worker_count = self.settings.general.worker_count
def process_one(scanned: ScannedFile) -> Tuple[ScannedFile, ClassificationResult]:
tokens = self.tokenizer.tokenize(scanned.path)
meta = self.analyzer.analyze(scanned.path)
classification = self.classifier.classify(scanned, tokens, meta)
# Match against known library shows if category is unsure or confidence is low
if known_shows and (
classification.category not in ("tv", "movie")
or classification.needs_quarantine
or classification.confidence < 0.8
):
from .library import match_known_show
matched = match_known_show(scanned.path.name, known_shows)
if not matched and len(scanned.path.parts) > 1:
matched = match_known_show(scanned.path.parent.name, known_shows)
if matched:
classification.category = "tv"
if not classification.tokens.title or classification.tokens.title.lower() in ("episode", "unknown", ""):
classification.tokens.title = matched["title"]
classification.confidence = max(classification.confidence, 0.95)
classification.needs_quarantine = False
return scanned, classification
completed = 0
with concurrent.futures.ThreadPoolExecutor(max_workers=worker_count) as executor:
future_to_file = {executor.submit(process_one, sf): sf for sf in qualifying_files}
for future in concurrent.futures.as_completed(future_to_file):
try:
res = future.result()
results.append(res)
except Exception as e:
sf = future_to_file[future]
logger.error("Error analyzing file", path=str(sf.path), error=str(e))
completed += 1
if progress_callback:
progress_callback(completed, total_files, str(future_to_file[future].path.name))
return results
def build_plan(
self, analysis_results: List[Tuple[ScannedFile, ClassificationResult]]
) -> List[PlannedOperation]:
"""Convert classification results into concrete planned operations."""
primary_dest_map: Dict[Path, Path] = {}
plan: List[PlannedOperation] = []
# 1. First pass: non-sidecar primary media files
for scanned, cls_res in analysis_results:
if not scanned.is_sidecar:
dst = self.namer.generate_destination_path(cls_res)
primary_dest_map[scanned.path] = dst
plan.append(
PlannedOperation(
src=scanned.path,
dst=dst,
action=self.settings.general.action,
category=cls_res.category,
confidence=cls_res.confidence,
details=cls_res.signals,
quarantine=cls_res.needs_quarantine,
quarantine_reason=cls_res.quarantine_reason,
)
)
# 2. Second pass: sidecars (subtitles, artwork, metadata) matching primary destinations
for scanned, cls_res in analysis_results:
if scanned.is_sidecar:
primary_dst = (
primary_dest_map.get(scanned.primary_media_path)
if scanned.primary_media_path
else None
)
dst = self.namer.generate_destination_path(cls_res, primary_dst_path=primary_dst)
plan.append(
PlannedOperation(
src=scanned.path,
dst=dst,
action=self.settings.general.action,
category=cls_res.category,
confidence=cls_res.confidence,
details=cls_res.signals,
quarantine=cls_res.needs_quarantine,
quarantine_reason=cls_res.quarantine_reason,
)
)
return plan
def run(
self,
dry_run: Optional[bool] = None,
progress_callback: Optional[Callable[[int, int, str], None]] = None,
filter_paths: Optional[List[Path]] = None,
show_name_override: Optional[str] = None,
) -> BatchExecutionReport:
"""Run full media sorter pipeline: scan, analyze, plan, and execute."""
start_time = time.time()
is_dry_run = self.settings.general.dry_run if dry_run is None else dry_run
logger.info("Executing media-sorter run", dry_run=is_dry_run)
# Discover and analyze
analysis_results = self.scan_and_analyze(
progress_callback=progress_callback, filter_paths=filter_paths
)
if show_name_override:
for scanned, cls_res in analysis_results:
cls_res.category = "tv"
cls_res.tokens.title = show_name_override
cls_res.tokens.is_episodic = True
cls_res.needs_quarantine = False
cls_res.confidence = max(cls_res.confidence, 0.95)
# Build plan
planned_ops = self.build_plan(analysis_results)
# Execute under inter-process lock to coordinate workers
lock_path = self.settings.get_database_path().with_suffix(".lock")
with acquire_process_lock(lock_path):
with get_db_session(self.engine) as session:
executor = MediaExecutor(self.settings, session)
# Check for crash recovery from prior runs
executor.recover_interrupted_batches()
report = executor.execute_batch(
planned_ops, dry_run=is_dry_run, progress_callback=progress_callback
)
elapsed = round(time.time() - start_time, 2)
logger.info(
"Media sorter run finished",
elapsed_seconds=elapsed,
total=report.total_files,
moved=report.moved_files,
quarantined=report.quarantined_files,
skipped=report.skipped_files,
failed=report.failed_files,
)
# Update library catalog with moved items
if not is_dry_run and report.moved_files > 0:
try:
from .library import record_detected_item
shows_dir = self.settings.get_destination_path("tv")
movies_dir = self.settings.get_destination_path("movie")
with get_db_session(self.engine) as session:
for op in report.operations:
if op.status == OperationStatus.COMMITTED.value:
cat = (op.category or "tv").lower()
dst_p = Path(op.dst)
if cat in ("tv", "anime"):
try:
rel_tv = dst_p.relative_to(shows_dir)
show_title = rel_tv.parts[0]
dest_folder = str(shows_dir / show_title)
record_detected_item(
session,
self.settings,
show_title,
"tv",
destination_folder=dest_folder,
delta_count=1,
)
except Exception:
pass
elif cat == "movie":
try:
rel_mv = dst_p.relative_to(movies_dir)
movie_title = rel_mv.parts[0]
dest_folder = str(movies_dir / movie_title)
record_detected_item(
session,
self.settings,
movie_title,
"movie",
destination_folder=dest_folder,
delta_count=1,
)
except Exception:
pass
except Exception as e:
logger.error("Error updating library from execution report", error=str(e))
# Send notification webhook if configured
if self.settings.notifications.enabled:
send_batch_notification(self.settings.notifications, report)
return report
def rollback(self, batch_id: Optional[str] = None) -> int:
"""Rollback a past batch of operations."""
with get_db_session(self.engine) as session:
executor = MediaExecutor(self.settings, session)
return executor.rollback_batch(batch_id)
def rollback_all(self) -> int:
"""Rollback all past completed batches."""
with get_db_session(self.engine) as session:
executor = MediaExecutor(self.settings, session)
return executor.rollback_all()
+271
View File
@@ -0,0 +1,271 @@
/* Base CSS for Media Sorter UI */
:root {
--bg: #ffffff; /* background */
--card-bg: #f9f9f9; /* cards */
--card-hover: #eaeaea;
--border: #dddddd;
--text: #222222;
--text-muted: #555555;
--accent: #0066ff; /* default accent */
--accent-hover: #0044cc;
--radius-sm: 4px;
--radius-md: 8px;
}
[data-theme="minimalistic"] {
/* Light, clean look */
--bg: #fafafa;
--card-bg: #ffffff;
--card-hover: #f0f0f0;
--border: #e0e0e0;
--text: #111111;
--text-muted: #777777;
--accent: #0077c2;
--accent-hover: #005599;
}
[data-theme="high-visibility"] {
/* Dark with bright accent for strong contrast */
--bg: #111111;
--card-bg: #1a1a1a;
--card-hover: #262626;
--border: #333333;
--text: #eeeeee;
--text-muted: #bbbbbb;
--accent: #ffdd00; /* bright yellow */
--accent-hover: #ffbb00;
}
[data-theme="pastel"] {
--bg: #fff8f0;
--card-bg: #ffffff;
--card-hover: #f0e6e0;
--border: #e6d8d1;
--text: #453636;
--text-muted: #7a5d5d;
--accent: #ff8c94; /* soft pink */
--accent-hover: #ff6b78;
}
[data-theme="grayscale"] {
--bg: #f5f5f5;
--card-bg: #ffffff;
--card-hover: #e0e0e0;
--border: #c0c0c0;
--text: #333333;
--text-muted: #777777;
--accent: #555555;
--accent-hover: #444444;
}
[data-theme="cyber"] {
/* Retain original cyber‑dark palette but with new ID */
--bg: #090d16;
--card-bg: #0f172a;
--card-hover: #1e293b;
--border: #334155;
--text: #f8fafc;
--text-muted: #94a3b8;
--accent: #38bdf8;
--accent-hover: #0284c7;
}
[data-theme="bios-amber"] {
--bg: #0c0800;
--card-bg: #181100;
--card-hover: #261b02;
--border: #4d3800;
--text: #ffb833;
--text-muted: #b37e1a;
--accent: #ff9900;
--accent-hover: #ffad33;
}
[data-theme="vapor-glitch"] {
--bg: #080312;
--card-bg: #130a24;
--card-hover: #1f113a;
--border: #38195a;
--text: #00f5d4;
--text-muted: #b388ff;
--accent: #f72585;
--accent-hover: #b5179e;
}
[data-theme="mossy-stone"] {
--bg: #111813;
--card-bg: #19241c;
--card-hover: #223227;
--border: #2d4234;
--text: #d2e0d5;
--text-muted: #859e8b;
--accent: #52b788;
--accent-hover: #40916c;
}
[data-theme="crimson-eclipse"] {
--bg: #0d0608;
--card-bg: #190c10;
--card-hover: #261217;
--border: #441822;
--text: #e8d0d5;
--text-muted: #a37581;
--accent: #ff4d6d;
--accent-hover: #c9184a;
}
[data-theme="blueprint-draft"] {
--bg: #0b1d3a;
--card-bg: #102a54;
--card-hover: #173b75;
--border: #1f4a91;
--text: #e2edfd;
--text-muted: #8db5e6;
--accent: #60a5fa;
--accent-hover: #3b82f6;
}
/* RGB mode styling */
.rgb-mode {
--rgb-duration: 5s;
}
.rgb-mode .card,
.rgb-mode .btn,
.rgb-mode .theme-select {
animation: rgbPulse var(--rgb-duration) infinite;
}
@keyframes rgbPulse {
0% { box-shadow: 0 0 8px rgba(255,0,0,0.5); }
33% { box-shadow: 0 0 8px rgba(0,255,0,0.5); }
66% { box-shadow: 0 0 8px rgba(0,0,255,0.5); }
100% { box-shadow: 0 0 8px rgba(255,0,0,0.5); }
}
[data-theme="neon-forest"] {
--bg: #001408;
--card-bg: #032410;
--card-hover: #07381b;
--border: #0d542a;
--text: #c8facc;
--text-muted: #5ea874;
--accent: #39ff14;
--accent-hover: #2ecc11;
}
[data-theme="retro-retro"] {
--bg: #120024;
--card-bg: #220038;
--card-hover: #330052;
--border: #6b0099;
--text: #fce7f3;
--text-muted: #d946ef;
--accent: #ff00ff;
--accent-hover: #d500d5;
}
[data-theme="golden-sand"] {
--bg: #1c150c;
--card-bg: #291e10;
--card-hover: #3b2c17;
--border: #594322;
--text: #fbf0dc;
--text-muted: #bda27e;
--accent: #ffb300;
--accent-hover: #e09d00;
}
[data-theme="deep-space"] {
--bg: #070913;
--card-bg: #0d1224;
--card-hover: #141c38;
--border: #232f57;
--text: #e2edfd;
--text-muted: #818cf8;
--accent: #6366f1;
--accent-hover: #4f46e5;
}
[data-theme="candy-cotton"] {
--bg: #1a1520;
--card-bg: #261f30;
--card-hover: #362c44;
--border: #524166;
--text: #fdf2f8;
--text-muted: #f472b6;
--accent: #ffb6c1;
--accent-hover: #f694a5;
}
/* General layout */
body {
background: var(--bg);
color: var(--text);
font-family: Arial, Helvetica, sans-serif;
margin: 0;
padding: 0;
}
header.app-header, footer.app-footer {
background: var(--card-bg);
border-bottom: 1px solid var(--border);
padding: 0.75rem 1rem;
display: flex;
justify-content: space-between;
align-items: center;
}
main.app-main {
display: flex;
flex-wrap: wrap;
gap: 1rem;
padding: 1rem;
}
.panel {
background: var(--card-bg);
border: 1px solid var(--border);
border-radius: var(--radius-md);
padding: 0.75rem;
flex: 1 1 300px;
max-width: 100%;
overflow: auto;
}
.btn {
background: var(--accent);
color: #fff;
border: none;
border-radius: var(--radius-sm);
padding: 0.4rem 0.8rem;
cursor: pointer;
}
.btn:hover {
background: var(--accent-hover);
}
.modal {
position: fixed;
inset: 0;
background: rgba(0,0,0,0.4);
display: flex;
align-items: center;
justify-content: center;
}
.modal.hidden { display: none; }
.modal-content {
background: var(--card-bg);
padding: 1.5rem;
border-radius: var(--radius-md);
min-width: 260px;
}
/* Compact mode tweaks */
[data-compact="true"] header.app-header, [data-compact="true"] footer.app-footer {
padding: 0.4rem 0.8rem;
font-size: 0.9rem;
}
[data-compact="true"] .panel {
padding: 0.5rem;
font-size: 0.85rem;
}
+36
View File
@@ -0,0 +1,36 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>Media Sorter Management</title>
<link rel="stylesheet" href="/static/css/base.css" />
<script src="/static/js/app.js" defer></script>
</head>
<body>
<!-- Settings Modal -->
<div id="settings-modal" class="modal hidden">
<div class="modal-content">
<h2>Settings</h2>
<label for="theme-select">Theme</label>
<select id="theme-select"></select>
<label for="compact-toggle"><input type="checkbox" id="compact-toggle" /> Compact Mode</label>
<button id="close-settings" class="btn">Close</button>
</div>
</div>
<!-- Main UI -->
<header class="app-header">
<h1>Media Sorter</h1>
<button id="open-settings" class="btn">Settings</button>
</header>
<main class="app-main">
<section id="folder-explorer" class="panel"></section>
<section id="library" class="panel"></section>
<section id="quarantine" class="panel"></section>
</main>
<footer class="app-footer">
<span id="status-bar">Ready</span>
</footer>
</body>
</html>
+653
View File
@@ -0,0 +1,653 @@
"""Filename and path tokenization engine for Media Sorter.
Robustly extracts semantic media tokens (title, year, season, episode, artist,
album, track, disc, quality, codec, release group, date stamps) from messy filenames,
scene releases, anime fansub conventions, and folder hierarchies.
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
from pathlib import Path
from typing import Dict, List, Optional, Tuple
# Regex patterns for Video & Episodic Media
ROMAN_NUMERALS: Dict[str, int] = {
"i": 1, "ii": 2, "iii": 3, "iv": 4, "v": 5,
"vi": 6, "vii": 7, "viii": 8, "ix": 9, "x": 10,
"xi": 11, "xii": 12, "xiii": 13, "xiv": 14, "xv": 15,
"xvi": 16, "xvii": 17, "xviii": 18, "xix": 19, "xx": 20,
}
RE_ROMAN_SEASON_EPISODE = re.compile(
r"""(?ix)
\bseason[\.\s_-]*(?P<season_roman>[ivx]+)[\.\s_-]*(?:episode|ep)[\.\s_-]*(?P<episode_roman>[ivx]+)\b
"""
)
RE_SEASON_EPISODE = re.compile(
r"""(?ix)
(?:
# Standard S01E02, S01E01-E02, S01E01E02, S01E01-02, S01E1171
(?<![0-9a-z])s(?P<season>\d{1,2})[\.\s_-]*(?:e|ep|ed|op)(?P<episode>\d{1,4})
(?:[\.\s_-]*(?:e|x|-|ep)(?P<episode_end>\d{1,4}))?(?![0-9])
|
# Scene 1x02, 2x01-02, 2x01-x02, 1x1171
(?<![0-9a-z])(?P<season_x>\d{1,2})x(?!(?:264|265|vid|hevc|avc))(?P<episode_x>\d{1,4})
(?:[\.\s_-]*(?:x|-)(?P<episode_x_end>\d{1,4}))?(?![0-9])
|
# Word season / episode: Season 1 Episode 2, Season 1 Episode 1171
\bseason[\.\s_-]*(?P<season_word>\d{1,2})[\.\s_-]*(?:episode|ep)[\.\s_-]*(?P<episode_word>\d{1,4})
(?:[\.\s_-]*(?:-|to)[\.\s_-]*(?:episode|ep)?[\\.\s_-]*(?P<episode_word_end>\d{1,4}))?\b
|
# Standalone episode: Episode 207, Ep 01, E233, E1171
(?<![0-9a-z])(?:episodes?|ep|e)[\.\s_-]*(?P<episode_standalone>\d{1,4})(?![0-9a-zA-Z])
)
"""
)
RE_SEASON_PACK = re.compile(
r"""(?ix)
(?<![0-9a-z])
(?:
s(?P<season_pack>\d{1,2})
|
season[\.\s_-]*(?P<season_pack_word>\d{1,2})
)
[\.\s_-]*(?:complete|full|season\.pack)\b
"""
)
# Anime fansub format: [ReleaseGroup] Show Title - 01 (or 01-02, or 01v2) - Optional Episode Title [1080p] [CRC32].mkv
RE_ANIME_RELEASE = re.compile(
r"""(?ix)
^\s*(?:\[(?P<group>[^\]]+)\][\s_]*)?
(?P<title>.+?)\s*(?:-\s*|_\-_\s*|_)\s*
(?P<episode>\d{1,4})(?:-(?P<episode_end>\d{1,4}))?(?:v\d+)?(?![0-9a-zA-Z])\s*
(?:\s*-\s*(?P<ep_title>[^\[\(]+?))?
(?:[\s_]*(?:\[?[0-9A-Fa-f]{8}\]?|\[(?P<tag>[^\]]+)\]|\((?P<tag_paren>[^\)]+)\))|[\s_]+[A-Za-z0-9_.-]+)*\s*\]?$
"""
)
# Space-separated anime format: [ReleaseGroup] Show Title 01 (Tags) [CRC32].mkv
RE_ANIME_RELEASE_SPACE = re.compile(
r"""(?ix)
^\s*\[(?P<group>[^\]]+)\]\s*
(?P<title>[^\[\(]+?)\s+
(?P<episode>\d{1,4})(?:-(?P<episode_end>\d{1,4}))?(?:v\d+)?(?![0-9a-zA-Z])\s*
(?:[\s_]*(?:\[?[0-9A-Fa-f]{8}\]?|\[(?P<tag>[^\]]+)\]|\((?P<tag_paren>[^\)]+)\))|[\s_]+[A-Za-z0-9_.-]+)*\s*\]?$
"""
)
# Underscore anime format: [ReleaseGroup]_Show_Title_01_[Tags].mp4
RE_ANIME_RELEASE_UNDERSCORE = re.compile(
r"""(?ix)
^\s*\[(?P<group>[^\]]+)\]_
(?P<title>[^\[\(]+?)_
(?P<episode>\d{1,4})(?:-(?P<episode_end>\d{1,4}))?(?:v\d+)?(?![0-9a-zA-Z])
(?:_*(?:\[?[0-9A-Fa-f]{8}\]?|\[(?P<tag>[^\]]+)\]|\((?P<tag_paren>[^\)]+)\))|_[A-Za-z0-9_.-]+)*\s*\]?$
"""
)
RE_ANIME_MOVIE = re.compile(
r"""(?ix)
^\s*\[(?P<group>[^\]]+)\]\s*
(?P<title>[^\[]+?)\s*
(?:\[(?P<tag>[^\]]+)\]|\((?P<tag_paren>[^\)]+)\))
"""
)
RE_YEAR_BOUND = re.compile(r"(?<![0-9a-zA-Z])(19\d{2}|20\d{2})(?![0-9a-zA-Z])")
RE_YEAR = re.compile(r"\b(19\d{2}|20\d{2})\b")
RE_EDITION = re.compile(
r"""(?ix)
\b(?P<edition>
directors?\.cut|director's\.cut|director's\scut
|
extended(?:\.cut|\.edition)?
|
remastered(?:\.edition)?|remaster
|
criterion(?:\.collection)?
|
final\.cut
|
theatrical(?:\.cut|\.version)?
|
unrated
|
special\.edition
|
imax(?:\.edition)?
|
ultimate\.edition
)\b
"""
)
EDITION_CANONICAL_MAP: Dict[str, str] = {
"extended": "Extended",
"extended.cut": "Extended",
"extended.edition": "Extended",
"directors.cut": "Director's Cut",
"director's.cut": "Director's Cut",
"director's cut": "Director's Cut",
"remastered": "Remastered",
"remastered.edition": "Remastered",
"remaster": "Remastered",
"criterion": "Criterion",
"criterion.collection": "Criterion",
"final.cut": "Final Cut",
"theatrical": "Theatrical",
"theatrical.cut": "Theatrical",
"theatrical.version": "Theatrical",
"unrated": "Unrated",
"special.edition": "Special Edition",
"imax": "IMAX",
"imax.edition": "IMAX",
"ultimate.edition": "Ultimate Edition",
}
RE_MOVIE_PART = re.compile(
r"""(?ix)
\b(?:cd|part|pt|disc)[\.\s_-]*(?P<part_num>\d{1,2})\b
"""
)
RE_DAILY_DATE = re.compile(
r"""(?ix)
(?<!\d)
(?P<year>19\d{2}|20\d{2})[-._]
(?P<month>0[1-9]|1[0-2])[-._]
(?P<day>0[1-9]|[12]\d|3[01])
(?!\d)
"""
)
# Technical specs
RE_RESOLUTION = re.compile(r"\b(2160p|4k|1080p|1080i|720p|576p|480p)\b", re.IGNORECASE)
RE_DIMENSIONS = re.compile(r"\b(?:\d{3,4})x(?P<height>2160|1080|720|576|480)\b", re.IGNORECASE)
RE_SOURCE = re.compile(r"\b(bluray|blu-ray|bdrip|web-dl|webrip|web|hdtv|dvdrip|dvd|remux)\b", re.IGNORECASE)
RE_VIDEO_CODEC = re.compile(r"\b(x265|x264|h\.?265|h\.?264|hevc|avc|av1|xvid|divx)\b", re.IGNORECASE)
RE_AUDIO_CODEC = re.compile(r"\b(truehd|atmos|dts-hd|dts|flac|aac|ac3|ddp?5\.1|mp3)\b", re.IGNORECASE)
RE_RELEASE_GROUP = re.compile(r"-([A-Za-z0-9_]+)(?:\[.*?\])?$", re.IGNORECASE)
RE_RELEASE_GROUP_UPGRADED = re.compile(
r"-(?:\[(?P<grp_bracket>[A-Za-z0-9_.-]+)\]|(?P<grp_plain>[A-Za-z0-9_]+))(?:\[.*?\])?$",
re.IGNORECASE,
)
RE_ILLEGAL_CHARS = re.compile(r'[<>:"/\\|?*\x00-\x1f]')
RE_TECH_ALL = re.compile(
r"""(?ix)
\b(
2160p|4k|1080p|1080i|720p|576p|480p
|
\d{3,4}x(?:2160|1080|720|576|480)
|
bluray|blu-ray|bdrip|web-dl|webrip|web|hdtv|dvdrip|dvd|remux
|
x265|x264|h\.?265|h\.?264|hevc|avc|av1|xvid|divx
|
truehd|atmos|dts-hd|dts|flac|aac|ac3|ddp?5\.1|mp3
|
directors?\.cut|director's\.cut|director's\scut|extended|remastered|criterion|final\.cut
|
cd\d|part\d|pt\d
|
proper
)\b
"""
)
KNOWN_ANIME_GROUPS = {
"subsplease", "horriblesubs", "erai-raws", "taigasubs", "judas", "commie", "asenshi", "coalgirls"
}
KNOWN_ANIME_TITLES = {
"naruto", "bleach", "one piece", "frieren", "dungeon meshi", "attack on titan",
"jujutsu kaisen", "mushoku tensei", "fairy tail", "fate stay night", "sword art online"
}
WINDOWS_RESERVED = {
"CON", "PRN", "AUX", "NUL",
"COM1", "COM2", "COM3", "COM4", "COM5", "COM6", "COM7", "COM8", "COM9",
"LPT1", "LPT2", "LPT3", "LPT4", "LPT5", "LPT6", "LPT7", "LPT8", "LPT9",
}
# Music / Audio track patterns: 01 - Title, 1-01 Title, Artist - 01 - Title
RE_MUSIC_TRACK = re.compile(
r"""(?ix)
^(?:(?P<disc>\d{1,2})[-_.])?(?P<track>\d{1,3})[\.\s_-]+(?P<title>.+)$
"""
)
# Photo and Home Video date stamps: IMG_20240812_142010, VID_20240812_142010, 2024-08-12 14.20.10
RE_CAMERA_DATE = re.compile(
r"""(?ix)
(?:img|vid|dsc|pano|mov)?[-_]?(?P<year>19\d{2}|20\d{2})[-_]?(?P<month>\d{2})[-_]?(?P<day>\d{2})
(?:[-_](?P<hour>\d{2})[-_]?(?P<minute>\d{2})[-_]?(?P<second>\d{2}))?
"""
)
# Podcast dated format: Show Name - 2026-03-15 - Episode Title
RE_PODCAST_DATE = re.compile(
r"""(?ix)
^(?P<show>.+?)\s*-\s*(?P<year>20\d{2})-(?P<month>\d{2})-(?P<day>\d{2})\s*-\s*(?P<title>.+)$
"""
)
# CRC32 checksum tag e.g. [194B3FBA]
RE_CRC32_TAG = re.compile(r"\[([0-9A-Fa-f]{8})\]")
@dataclass
class TokenizedFilename:
raw_name: str
title: Optional[str] = None
year: Optional[int] = None
season: Optional[int] = None
episode: Optional[int] = None
multi_episodes: List[int] = field(default_factory=list)
episode_title: Optional[str] = None
artist: Optional[str] = None
album: Optional[str] = None
track: Optional[int] = None
disc: Optional[int] = None
group: Optional[str] = None
resolution: Optional[str] = None
source: Optional[str] = None
video_codec: Optional[str] = None
audio_codec: Optional[str] = None
crc32: Optional[str] = None
date_stamp: Optional[str] = None
air_date: Optional[str] = None
edition: Optional[str] = None
part: Optional[int] = None
part_label: Optional[str] = None
is_anime: bool = False
is_episodic: bool = False
is_music: bool = False
is_photo_or_home_video: bool = False
is_daily: bool = False
is_season_pack: bool = False
# Contract alias per PROJECT.md:74
TokenizedMedia = TokenizedFilename
class FilenameTokenizer:
"""Parses raw filenames and directory paths into semantic tokens."""
def tokenize(self, file_path: Path) -> TokenizedFilename:
stem = file_path.stem
raw_name = file_path.name
tokens = TokenizedFilename(raw_name=raw_name)
ext = file_path.suffix.lower()
# Check for CRC32 tag
crc_m = RE_CRC32_TAG.search(stem)
if crc_m:
tokens.crc32 = crc_m.group(1).upper()
# Check Windows reserved names
if stem.upper() in WINDOWS_RESERVED:
tokens.title = stem.upper()
tokens.is_photo_or_home_video = True
return tokens
# Pre-clean illegal characters
had_illegal = False
if RE_ILLEGAL_CHARS.search(stem):
had_illegal = True
stem = RE_ILLEGAL_CHARS.sub(" ", stem)
stem = re.sub(r"\s+", " ", stem).strip()
# 1. Technical specifications
res_m = RE_RESOLUTION.search(stem)
if res_m:
tokens.resolution = res_m.group(1).lower()
if tokens.resolution == "4k":
tokens.resolution = "2160p"
else:
dim_m = RE_DIMENSIONS.search(stem)
if dim_m:
tokens.resolution = f"{dim_m.group('height')}p"
src_m = RE_SOURCE.search(stem)
if src_m:
tokens.source = src_m.group(1).upper()
vc_m = RE_VIDEO_CODEC.search(stem)
if vc_m:
tokens.video_codec = vc_m.group(1).lower().replace(".", "")
ac_m = RE_AUDIO_CODEC.search(stem)
if ac_m:
tokens.audio_codec = ac_m.group(1).upper()
# Check edition
ed_m = RE_EDITION.search(stem)
if ed_m:
ed_raw = ed_m.group("edition").lower().replace(" ", ".")
tokens.edition = EDITION_CANONICAL_MAP.get(ed_raw, ed_m.group("edition"))
# Check part
pt_m = RE_MOVIE_PART.search(stem)
if pt_m:
tokens.part = int(pt_m.group("part_num"))
tokens.part_label = f"Pt.{tokens.part}"
# 2. Check for Podcast date format
pod_m = RE_PODCAST_DATE.match(stem)
if pod_m:
tokens.title = pod_m.group("title").strip()
tokens.artist = pod_m.group("show").strip()
tokens.year = int(pod_m.group("year"))
tokens.date_stamp = f"{pod_m.group('year')}-{pod_m.group('month')}-{pod_m.group('day')}"
tokens.air_date = tokens.date_stamp
return tokens
# 3. Check for Daily / Broadcast dated format (TV or Podcast)
daily_m = RE_DAILY_DATE.search(stem)
if daily_m:
y, m, d = daily_m.group("year"), daily_m.group("month"), daily_m.group("day")
date_str = f"{y}-{m}-{d}"
tokens.date_stamp = date_str
tokens.air_date = date_str
tokens.year = int(y)
prefix = stem[: daily_m.start()]
clean_pfx = self._clean_title(prefix)
tokens.title = clean_pfx
if ext in {".mp3", ".flac", ".ogg", ".m4a", ".aac"}:
tokens.artist = clean_pfx
else:
tokens.is_daily = True
tokens.is_episodic = True
tokens.season = int(y)
return tokens
# 4. Check for Camera / Date stamp (Photos & Home Videos)
cam_m = RE_CAMERA_DATE.search(stem)
if cam_m:
y, m, d = cam_m.group("year"), cam_m.group("month"), cam_m.group("day")
tokens.date_stamp = f"{y}-{m}-{d}"
tokens.year = int(y)
tokens.is_photo_or_home_video = True
return tokens
# 5. Check Roman Numeral TV pattern: Rome.Season.II.Episode.IV
roman_m = RE_ROMAN_SEASON_EPISODE.search(stem)
if roman_m:
tokens.is_episodic = True
s_rom = roman_m.group("season_roman").lower()
e_rom = roman_m.group("episode_roman").lower()
tokens.season = ROMAN_NUMERALS.get(s_rom, 1)
tokens.episode = ROMAN_NUMERALS.get(e_rom, 1)
prefix = stem[: roman_m.start()]
tokens.title = self._clean_title(prefix)
return tokens
# 6. Check TV Season Pack: Succession.S02.Complete
pack_m = RE_SEASON_PACK.search(stem)
if pack_m:
tokens.is_episodic = True
tokens.is_season_pack = True
s_val = pack_m.group("season_pack") or pack_m.group("season_pack_word")
tokens.season = int(s_val)
prefix = stem[: pack_m.start()]
tokens.title = self._clean_title(prefix)
return tokens
# 7. Check Standard TV episodic patterns (S01E02, 1x02, Season 1 Episode 2, Episode 207)
tv_m = RE_SEASON_EPISODE.search(stem)
if tv_m:
tokens.is_episodic = True
season_str = tv_m.group("season") or tv_m.group("season_x") or tv_m.group("season_word")
ep_str = tv_m.group("episode") or tv_m.group("episode_x") or tv_m.group("episode_word") or tv_m.group("episode_standalone")
if season_str:
tokens.season = int(season_str)
else:
tokens.season = self._extract_season_from_path(file_path) or 1
if ep_str:
tokens.episode = int(ep_str)
end_ep = tv_m.group("episode_end") or tv_m.group("episode_x_end") or tv_m.group("episode_word_end")
if end_ep:
tokens.multi_episodes = list(range(tokens.episode, int(end_ep) + 1))
# If filename had illegal characters and matched standalone episode (e.g. Show: "Special" <Episode> | 1?.mkv)
if had_illegal and tv_m.group("episode_standalone"):
tokens.title = self._clean_title(stem)
return tokens
# Extract title before season marker
prefix = stem[: tv_m.start()]
clean_pfx = self._clean_title(prefix)
if clean_pfx:
tokens.title = clean_pfx
else:
tokens.title = self._extract_title_from_context(file_path) or "Episode"
# Check for year in prefix using RE_YEAR_BOUND
if prefix:
yr_m = RE_YEAR_BOUND.search(prefix)
if yr_m:
tokens.year = int(yr_m.group(1))
tokens.title = self._clean_title(prefix[: yr_m.start()])
# Extract episode title after season marker
suffix = stem[tv_m.end() :]
ep_title = self._extract_episode_title(suffix)
if ep_title:
tokens.episode_title = ep_title
# Release group at end
grp_m = RE_RELEASE_GROUP_UPGRADED.search(stem)
if grp_m:
tokens.group = grp_m.group("grp_bracket") or grp_m.group("grp_plain")
# Check if title is a known anime title
if tokens.title and tokens.title.lower() in KNOWN_ANIME_TITLES:
tokens.is_anime = True
return tokens
# 8. Check Anime fansub format: [Group] Title - 01 - Episode Title [1080p]
anime_m = (
RE_ANIME_RELEASE.match(stem)
or RE_ANIME_RELEASE_SPACE.match(stem)
or RE_ANIME_RELEASE_UNDERSCORE.match(stem)
)
if anime_m and (anime_m.group("group") or ext in {".mkv", ".mp4", ".avi", ".mov", ".ts", ".webm", ".m4v", ".flv"}):
ep_val = int(anime_m.group("episode"))
grp_name = anime_m.group("group").strip() if anime_m.group("group") else None
raw_title = anime_m.group("title")
title_clean = self._clean_title(raw_title, preserve_paren=True)
# Distinguish movie year from anime episode
is_known_anime = title_clean.lower() in KNOWN_ANIME_TITLES or (grp_name and grp_name.lower() in KNOWN_ANIME_GROUPS)
has_explicit_season = self._extract_season_from_path(file_path) is not None
has_range = bool(anime_m.groupdict().get("episode_end"))
if 1900 <= ep_val <= 2099 and not has_range and not is_known_anime and not has_explicit_season:
tokens.year = ep_val
tokens.title = title_clean
tokens.group = grp_name
tokens.is_anime = False
tokens.is_episodic = False
return tokens
else:
tokens.is_anime = True
tokens.group = grp_name
tokens.title = title_clean
tokens.episode = ep_val
if anime_m.groupdict().get("ep_title"):
tokens.episode_title = self._clean_title(anime_m.group("ep_title"))
if anime_m.groupdict().get("episode_end"):
tokens.multi_episodes = list(range(ep_val, int(anime_m.group("episode_end")) + 1))
tokens.season = self._extract_season_from_path(file_path) or 1
tokens.is_episodic = True
return tokens
# 9. Check Anime movie format: [Judas] Fate Stay Night... [BD 1080p]
anime_mov_m = RE_ANIME_MOVIE.match(stem)
if anime_mov_m:
grp = anime_mov_m.group("group").strip()
if grp.lower() in KNOWN_ANIME_GROUPS:
tokens.group = grp
tokens.is_anime = True
tokens.title = self._clean_title(anime_mov_m.group("title"), preserve_dots=True)
return tokens
# 10. Check for Music track pattern
mus_m = RE_MUSIC_TRACK.match(stem)
if mus_m:
tokens.is_music = True
tokens.track = int(mus_m.group("track"))
if mus_m.group("disc"):
tokens.disc = int(mus_m.group("disc"))
tokens.title = self._clean_title(mus_m.group("title"))
parent = file_path.parent
if parent and parent.name:
parts = parent.name.split(" - ")
if len(parts) >= 2:
tokens.artist = parts[0].strip()
tokens.album = parts[1].strip()
return tokens
# 11. Movie pattern: Title (Year) or Title.Year.Quality
# Parenthesized year first
paren_yr = re.search(r"\((19\d{2}|20\d{2})\)", stem)
if paren_yr:
tokens.year = int(paren_yr.group(1))
prefix = stem[: paren_yr.start()]
tokens.title = self._clean_title(prefix)
grp_m = RE_RELEASE_GROUP_UPGRADED.search(stem)
if grp_m:
tokens.group = grp_m.group("grp_bracket") or grp_m.group("grp_plain")
return tokens
# Delimiter-based right-to-left year detection
tech_start = len(stem)
for m in RE_TECH_ALL.finditer(stem):
if m.start() < tech_start:
tech_start = m.start()
year_matches = list(RE_YEAR_BOUND.finditer(stem))
if year_matches:
valid_matches = [m for m in year_matches if m.start() <= tech_start]
if not valid_matches:
valid_matches = year_matches
best_match = valid_matches[-1]
tokens.year = int(best_match.group(1))
prefix = stem[: best_match.start()]
tokens.title = self._clean_title(prefix)
grp_m = RE_RELEASE_GROUP_UPGRADED.search(stem)
if grp_m:
tokens.group = grp_m.group("grp_bracket") or grp_m.group("grp_plain")
return tokens
# Fallback: strip tech specs and clean whole stem as title
prefix = stem[:tech_start].strip(" .-_")
tokens.title = self._clean_title(prefix if prefix else stem)
grp_m = RE_RELEASE_GROUP_UPGRADED.search(stem)
if grp_m:
tokens.group = grp_m.group("grp_bracket") or grp_m.group("grp_plain")
return tokens
def _clean_title(self, raw: str, preserve_paren: bool = False, preserve_dots: bool = False) -> str:
"""Replace dots, underscores, and scene separators with clean spaces."""
# Strip leading bracket tags like [YTS.MX] or [SubsPlease] if present
raw = re.sub(r"^\s*\[[^\]]+\]\s*", "", raw)
# Protect decimal numbers like 2.5, 3.5, 4.5, 1.5, 0.5, 1.11, etc.
raw = re.sub(r"(?<=\d)\.(?=\d)", "PROTECTEDDECIMALDOT", raw)
if preserve_dots:
cleaned = re.sub(r"_+", " ", raw).strip()
else:
# Replace dots with space, except if dot is followed by space in Roman numeral (e.g. "I. ")
cleaned = re.sub(r"(?<=\b[IVXLCDM])\.\s+", "._KEEP_DOT_SPACE_", raw)
cleaned = re.sub(r"[\._]+", " ", cleaned)
cleaned = cleaned.replace("._KEEP_DOT_SPACE_", ". ")
cleaned = cleaned.replace("PROTECTEDDECIMALDOT", ".")
cleaned = re.sub(r"\s+", " ", cleaned).strip()
if preserve_paren and re.search(r"\([12]\d{3}\)$", cleaned):
# Do not strip trailing parenthesis if it's (Year)
pass
else:
cleaned = re.sub(r"[\-\(\)\[\]]+$", "", cleaned).strip()
return cleaned
def _extract_episode_title(self, suffix: str) -> Optional[str]:
"""Extract episode title from string after SxxExx marker, stripping tech tags."""
s = suffix.strip(" .-_")
if not s:
return None
# Split by known tech tags
for reg in (RE_RESOLUTION, RE_SOURCE, RE_VIDEO_CODEC, RE_AUDIO_CODEC, RE_EDITION):
m = reg.search(s)
if m:
s = s[: m.start()].strip(" .-_")
# Strip release group
grp = RE_RELEASE_GROUP_UPGRADED.search(s)
if grp:
s = s[: grp.start()].strip(" .-_")
cleaned = self._clean_title(s)
return cleaned if cleaned else None
def _extract_season_from_path(self, file_path: Path) -> Optional[int]:
"""Attempt to extract season number from parent directory names like 'Season 2', 'Season 08 - Water Seven', or 'S03'."""
try:
system_folders = {
"downloads", "jdownloads", "media", "completed", "incomplete", "torrent",
"torrents", "root", "home", "mnt", "md0", "storage", "tmp", "temp", "var", "etc", "usr"
}
for part in reversed(file_path.parts[:-1]):
if part.lower().strip() in system_folders:
break
m = re.search(r"(?i)\b(?:season|series|s)\s*(\d{1,2})\b", part)
if m:
return int(m.group(1))
except Exception:
pass
return None
def _extract_title_from_context(self, file_path: Path) -> Optional[str]:
"""Attempt to extract show title from immediate parent directory (e.g. Show/Episode 01.mkv or Show/Season 1/Ep01.mkv)."""
try:
parent = file_path.parent
if not parent or str(parent) in ("/", ".", ""):
return None
p_name = parent.name
if re.search(r"(?i)\b(?:season|series|s)\s*\d+\b", p_name):
# Ascend one level if inside a season folder
parent = parent.parent
p_name = parent.name if parent else ""
if not p_name:
return None
p_lower = p_name.lower().strip()
system_folders = {
"downloads", "jdownloads", "media", "completed", "incomplete", "torrent",
"torrents", "root", "home", "mnt", "md0", "storage", "tmp", "temp", "var", "etc", "usr"
}
if p_lower in system_folders:
return None
clean = self._clean_title(p_name)
if len(clean) >= 2:
return clean
except Exception:
pass
return None
+5
View File
@@ -0,0 +1,5 @@
"""Media Sorter E2E Benchmark Test Suite & Offline Runner package."""
from .benchmark_cases import BENCHMARK_CASES, CASES_BY_DOMAIN, CASES_BY_ID, BenchmarkCase
__all__ = ["BenchmarkCase", "BENCHMARK_CASES", "CASES_BY_ID", "CASES_BY_DOMAIN"]
File diff suppressed because it is too large. Load diff
+899
View File
@@ -0,0 +1,899 @@
{
"metadata": {
"timestamp": "2026-09-06T02:50:04.839136+00:00",
"python_version": "3.14.4",
"total_duration_ms": 112.93,
"runner": "MediaSorter Offline E2E Benchmark Runner"
},
"summary": {
"total_cases": 64,
"passed_cases": 64,
"failed_cases": 0,
"pass_rate_pct": 100.0
},
"domains": {
"Standard TV": {
"total": 11,
"passed": 11,
"failed": 0,
"pass_rate_pct": 100.0,
"total_duration_ms": 19.434,
"avg_duration_ms": 1.77
},
"Anime": {
"total": 11,
"passed": 11,
"failed": 0,
"pass_rate_pct": 100.0,
"total_duration_ms": 18.826,
"avg_duration_ms": 1.71
},
"Movies": {
"total": 14,
"passed": 14,
"failed": 0,
"pass_rate_pct": 100.0,
"total_duration_ms": 24.04,
"avg_duration_ms": 1.72
},
"Specials & Extras": {
"total": 9,
"passed": 9,
"failed": 0,
"pass_rate_pct": 100.0,
"total_duration_ms": 17.233999999999998,
"avg_duration_ms": 1.91
},
"Daily / Dated Shows": {
"total": 9,
"passed": 9,
"failed": 0,
"pass_rate_pct": 100.0,
"total_duration_ms": 15.488000000000001,
"avg_duration_ms": 1.72
},
"Messy & Complex": {
"total": 10,
"passed": 10,
"failed": 0,
"pass_rate_pct": 100.0,
"total_duration_ms": 17.905,
"avg_duration_ms": 1.79
}
},
"failures": [],
"all_results": [
{
"id": "TV-01",
"domain": "Standard TV",
"filename": "Breaking.Bad.S05E14.Ozymandias.1080p.BluRay.x264-ROVERS.mkv",
"edge_case_type": "standard_sxxexx",
"passed": true,
"duration_ms": 2.445,
"diffs": {},
"actual_category": "tv",
"actual_title": "Breaking Bad",
"actual_destination_subpath": "TV Shows/Breaking Bad/Season 05/Breaking Bad - S05E14.mkv",
"exception": null
},
{
"id": "TV-02",
"domain": "Standard TV",
"filename": "Stranger.Things.S04E01-E02.Chapter.One.720p.WEB-DL.mkv",
"edge_case_type": "multi_episode_hyphen_range",
"passed": true,
"duration_ms": 1.765,
"diffs": {},
"actual_category": "tv",
"actual_title": "Stranger Things",
"actual_destination_subpath": "TV Shows/Stranger Things/Season 04/Stranger Things - S04E01-E02.mkv",
"exception": null
},
{
"id": "TV-03",
"domain": "Standard TV",
"filename": "House.M.D.S03E01E02.1080p.mkv",
"edge_case_type": "multi_episode_concatenated",
"passed": true,
"duration_ms": 1.716,
"diffs": {},
"actual_category": "tv",
"actual_title": "House M D",
"actual_destination_subpath": "TV Shows/House M D/Season 03/House M D - S03E01-E02.mkv",
"exception": null
},
{
"id": "TV-04",
"domain": "Standard TV",
"filename": "The.Wire.1x09.HDTV.mkv",
"edge_case_type": "scene_season_x_episode",
"passed": true,
"duration_ms": 1.695,
"diffs": {},
"actual_category": "tv",
"actual_title": "The Wire",
"actual_destination_subpath": "TV Shows/The Wire/Season 01/The Wire - S01E09.mkv",
"exception": null
},
{
"id": "TV-05",
"domain": "Standard TV",
"filename": "The.Office.2x01-02.mkv",
"edge_case_type": "scene_multi_episode_range",
"passed": true,
"duration_ms": 1.692,
"diffs": {},
"actual_category": "tv",
"actual_title": "The Office",
"actual_destination_subpath": "TV Shows/The Office/Season 02/The Office - S02E01-E02.mkv",
"exception": null
},
{
"id": "TV-06",
"domain": "Standard TV",
"filename": "Rome.Season.II.Episode.IV.mkv",
"edge_case_type": "roman_numerals",
"passed": true,
"duration_ms": 1.68,
"diffs": {},
"actual_category": "tv",
"actual_title": "Rome",
"actual_destination_subpath": "TV Shows/Rome/Season 02/Rome - S02E04.mkv",
"exception": null
},
{
"id": "TV-07",
"domain": "Standard TV",
"filename": "Doctor Who Season 5 Episode 1 Eleventh Hour.mkv",
"edge_case_type": "word_season_episode",
"passed": true,
"duration_ms": 1.705,
"diffs": {},
"actual_category": "tv",
"actual_title": "Doctor Who",
"actual_destination_subpath": "TV Shows/Doctor Who/Season 05/Doctor Who - S05E01.mkv",
"exception": null
},
{
"id": "TV-08",
"domain": "Standard TV",
"filename": "Succession.S02.Complete.1080p.WEB-DL.mkv",
"edge_case_type": "season_pack_complete",
"passed": true,
"duration_ms": 1.688,
"diffs": {},
"actual_category": "tv",
"actual_title": "Succession",
"actual_destination_subpath": "TV Shows/Succession/Season 02/Succession - Season 02.mkv",
"exception": null
},
{
"id": "TV-09",
"domain": "Standard TV",
"filename": "Game.of.Thrones.S08E03.720p.HDTV.x264-AVS.mkv",
"edge_case_type": "standard_sxxexx_scene_group",
"passed": true,
"duration_ms": 1.69,
"diffs": {},
"actual_category": "tv",
"actual_title": "Game of Thrones",
"actual_destination_subpath": "TV Shows/Game of Thrones/Season 08/Game of Thrones - S08E03.mkv",
"exception": null
},
{
"id": "TV-10",
"domain": "Standard TV",
"filename": "Chernobyl.S01E05.Vichnaya.Pamyat.1080p.mkv",
"edge_case_type": "sxxexx_with_episode_title",
"passed": true,
"duration_ms": 1.675,
"diffs": {},
"actual_category": "tv",
"actual_title": "Chernobyl",
"actual_destination_subpath": "TV Shows/Chernobyl/Season 01/Chernobyl - S01E05.mkv",
"exception": null
},
{
"id": "TV-11",
"domain": "Standard TV",
"filename": "Friends.S06E15-E16.The.One.That.Could.Have.Been.mkv",
"edge_case_type": "multi_episode_double_digit",
"passed": true,
"duration_ms": 1.683,
"diffs": {},
"actual_category": "tv",
"actual_title": "Friends",
"actual_destination_subpath": "TV Shows/Friends/Season 06/Friends - S06E15-E16.mkv",
"exception": null
},
{
"id": "ANIME-01",
"domain": "Anime",
"filename": "[SubsPlease] Frieren - Beyond Journey's End - 01 (1080p) [ABCD1234].mkv",
"edge_case_type": "fansub_brackets_crc",
"passed": true,
"duration_ms": 1.762,
"diffs": {},
"actual_category": "anime",
"actual_title": "Frieren - Beyond Journey's End",
"actual_destination_subpath": "Anime/Frieren - Beyond Journey's End/Frieren - Beyond Journey's End - 01 [SubsPlease].mkv",
"exception": null
},
{
"id": "ANIME-02",
"domain": "Anime",
"filename": "[SubsPlease] 葬送のフリーレン - 12 (1080p) [98E7B1A2].mkv",
"edge_case_type": "unicode_kanji_title",
"passed": true,
"duration_ms": 1.714,
"diffs": {},
"actual_category": "anime",
"actual_title": "葬送のフリーレン",
"actual_destination_subpath": "Anime/葬送のフリーレン/葬送のフリーレン - 12 [SubsPlease].mkv",
"exception": null
},
{
"id": "ANIME-03",
"domain": "Anime",
"filename": "[HorribleSubs] Fairy Tail (2014) - 176 [720p].mkv",
"edge_case_type": "parenthesized_title_year",
"passed": true,
"duration_ms": 1.688,
"diffs": {},
"actual_category": "anime",
"actual_title": "Fairy Tail (2014)",
"actual_destination_subpath": "Anime/Fairy Tail (2014)/Fairy Tail (2014) - 176 [HorribleSubs].mkv",
"exception": null
},
{
"id": "ANIME-04",
"domain": "Anime",
"filename": "[Erai-raws] One Piece - 1088 [1080p].mkv",
"edge_case_type": "four_digit_absolute_numbering",
"passed": true,
"duration_ms": 1.688,
"diffs": {},
"actual_category": "anime",
"actual_title": "One Piece",
"actual_destination_subpath": "Anime/One Piece/One Piece - 1088 [Erai-raws].mkv",
"exception": null
},
{
"id": "ANIME-05",
"domain": "Anime",
"filename": "[SubsPlease] Dungeon Meshi - 01-02 (1080p).mkv",
"edge_case_type": "multi_episode_anime",
"passed": true,
"duration_ms": 1.7,
"diffs": {},
"actual_category": "anime",
"actual_title": "Dungeon Meshi",
"actual_destination_subpath": "Anime/Dungeon Meshi/Dungeon Meshi - 01-02 [SubsPlease].mkv",
"exception": null
},
{
"id": "ANIME-06",
"domain": "Anime",
"filename": "[TaigaSubs] Attack on Titan OVA - 01 [720p].mkv",
"edge_case_type": "anime_ova",
"passed": true,
"duration_ms": 1.709,
"diffs": {},
"actual_category": "anime",
"actual_title": "Attack on Titan OVA",
"actual_destination_subpath": "Anime/Attack on Titan OVA/Attack on Titan OVA - 01 [TaigaSubs].mkv",
"exception": null
},
{
"id": "ANIME-07",
"domain": "Anime",
"filename": "Naruto Episode 207 The Supposed Sealed Ability.mkv",
"edge_case_type": "standalone_episode_keyword",
"passed": true,
"duration_ms": 1.715,
"diffs": {},
"actual_category": "anime",
"actual_title": "Naruto",
"actual_destination_subpath": "Anime/Naruto/Naruto - 207.mkv",
"exception": null
},
{
"id": "ANIME-08",
"domain": "Anime",
"filename": "[Judas] Fate Stay Night - Heaven's Feel - I. Presage Flower [BD 1080p].mkv",
"edge_case_type": "anime_movie_roman_numeral",
"passed": true,
"duration_ms": 1.754,
"diffs": {},
"actual_category": "anime",
"actual_title": "Fate Stay Night - Heaven's Feel - I. Presage Flower",
"actual_destination_subpath": "Anime/Fate Stay Night - Heaven's Feel - I. Presage Flower/Fate Stay Night - Heaven's Feel - I. Presage Flower [Judas].mkv",
"exception": null
},
{
"id": "ANIME-09",
"domain": "Anime",
"filename": "BLEACH - Sennen Kessen-hen - 27 [E89717B7].mkv",
"edge_case_type": "no_group_fansub_crc",
"passed": true,
"duration_ms": 1.677,
"diffs": {},
"actual_category": "anime",
"actual_title": "BLEACH - Sennen Kessen-hen",
"actual_destination_subpath": "Anime/BLEACH - Sennen Kessen-hen/BLEACH - Sennen Kessen-hen - 27.mkv",
"exception": null
},
{
"id": "ANIME-10",
"domain": "Anime",
"filename": "[Erai-raws] Jujutsu Kaisen 2nd Season - 14 [1080p][Multiple Subtitle].mkv",
"edge_case_type": "cour_season_title_tags",
"passed": true,
"duration_ms": 1.715,
"diffs": {},
"actual_category": "anime",
"actual_title": "Jujutsu Kaisen 2nd Season",
"actual_destination_subpath": "Anime/Jujutsu Kaisen 2nd Season/Jujutsu Kaisen 2nd Season - 14 [Erai-raws].mkv",
"exception": null
},
{
"id": "ANIME-11",
"domain": "Anime",
"filename": "[SubsPlease] Mushoku Tensei S2 - 18 (1080p) [F28B1452].mkv",
"edge_case_type": "short_season_notation",
"passed": true,
"duration_ms": 1.704,
"diffs": {},
"actual_category": "anime",
"actual_title": "Mushoku Tensei S2",
"actual_destination_subpath": "Anime/Mushoku Tensei S2/Mushoku Tensei S2 - 18 [SubsPlease].mkv",
"exception": null
},
{
"id": "MOVIE-01",
"domain": "Movies",
"filename": "Inception.2010.1080p.BluRay.x264-FraMeSToR.mkv",
"edge_case_type": "standard_movie_year",
"passed": true,
"duration_ms": 1.792,
"diffs": {},
"actual_category": "movie",
"actual_title": "Inception",
"actual_destination_subpath": "Movies/Inception (2010)/Inception (2010).mkv",
"exception": null
},
{
"id": "MOVIE-02",
"domain": "Movies",
"filename": "1917.2019.1080p.BluRay.x264.mkv",
"edge_case_type": "numerical_title_with_year",
"passed": true,
"duration_ms": 1.705,
"diffs": {},
"actual_category": "movie",
"actual_title": "1917",
"actual_destination_subpath": "Movies/1917 (2019)/1917 (2019).mkv",
"exception": null
},
{
"id": "MOVIE-03",
"domain": "Movies",
"filename": "2001.A.Space.Odyssey.1968.REMASTERED.1080p.mkv",
"edge_case_type": "numerical_title_year_and_edition",
"passed": true,
"duration_ms": 1.72,
"diffs": {},
"actual_category": "movie",
"actual_title": "2001 A Space Odyssey",
"actual_destination_subpath": "Movies/2001 A Space Odyssey (1968)/2001 A Space Odyssey (1968) [Remastered].mkv",
"exception": null
},
{
"id": "MOVIE-04",
"domain": "Movies",
"filename": "Blade.Runner.2049.2017.2160p.UHD.BluRay.x265.mkv",
"edge_case_type": "future_year_in_title",
"passed": true,
"duration_ms": 1.703,
"diffs": {},
"actual_category": "movie",
"actual_title": "Blade Runner 2049",
"actual_destination_subpath": "Movies/Blade Runner 2049 (2017)/Blade Runner 2049 (2017).mkv",
"exception": null
},
{
"id": "MOVIE-05",
"domain": "Movies",
"filename": "Wonder.Woman.1984.2020.1080p.WEB-DL.mkv",
"edge_case_type": "year_in_title",
"passed": true,
"duration_ms": 1.718,
"diffs": {},
"actual_category": "movie",
"actual_title": "Wonder Woman 1984",
"actual_destination_subpath": "Movies/Wonder Woman 1984 (2020)/Wonder Woman 1984 (2020).mkv",
"exception": null
},
{
"id": "MOVIE-06",
"domain": "Movies",
"filename": "Class.of.1999.1990.720p.mkv",
"edge_case_type": "past_year_in_title",
"passed": true,
"duration_ms": 1.703,
"diffs": {},
"actual_category": "movie",
"actual_title": "Class of 1999",
"actual_destination_subpath": "Movies/Class of 1999 (1990)/Class of 1999 (1990).mkv",
"exception": null
},
{
"id": "MOVIE-07",
"domain": "Movies",
"filename": "Titanic.1997.DVD.CD1.avi",
"edge_case_type": "multi_part_cd1",
"passed": true,
"duration_ms": 1.696,
"diffs": {},
"actual_category": "movie",
"actual_title": "Titanic",
"actual_destination_subpath": "Movies/Titanic (1997)/Titanic (1997) [Pt.1].avi",
"exception": null
},
{
"id": "MOVIE-08",
"domain": "Movies",
"filename": "Titanic.1997.DVD.CD2.avi",
"edge_case_type": "multi_part_cd2",
"passed": true,
"duration_ms": 1.692,
"diffs": {},
"actual_category": "movie",
"actual_title": "Titanic",
"actual_destination_subpath": "Movies/Titanic (1997)/Titanic (1997) [Pt.2].avi",
"exception": null
},
{
"id": "MOVIE-09",
"domain": "Movies",
"filename": "Kill.Bill.Vol.1.2003.1080p.BluRay.mkv",
"edge_case_type": "volume_in_movie_title",
"passed": true,
"duration_ms": 1.71,
"diffs": {},
"actual_category": "movie",
"actual_title": "Kill Bill Vol 1",
"actual_destination_subpath": "Movies/Kill Bill Vol 1 (2003)/Kill Bill Vol 1 (2003).mkv",
"exception": null
},
{
"id": "MOVIE-10",
"domain": "Movies",
"filename": "The.Lord.of.the.Rings.The.Fellowship.of.the.Ring.2001.Extended.1080p.mkv",
"edge_case_type": "edition_extended",
"passed": true,
"duration_ms": 1.713,
"diffs": {},
"actual_category": "movie",
"actual_title": "The Lord of the Rings The Fellowship of the Ring",
"actual_destination_subpath": "Movies/The Lord of the Rings The Fellowship of the Ring (2001)/The Lord of the Rings The Fellowship of the Ring (2001) [Extended].mkv",
"exception": null
},
{
"id": "MOVIE-11",
"domain": "Movies",
"filename": "Aliens.1986.Directors.Cut.1080p.BluRay.mkv",
"edge_case_type": "edition_directors_cut",
"passed": true,
"duration_ms": 1.689,
"diffs": {},
"actual_category": "movie",
"actual_title": "Aliens",
"actual_destination_subpath": "Movies/Aliens (1986)/Aliens (1986) [Director's Cut].mkv",
"exception": null
},
{
"id": "MOVIE-12",
"domain": "Movies",
"filename": "Gladiator.2000.Remastered.1080p.BluRay.mkv",
"edge_case_type": "edition_remastered",
"passed": true,
"duration_ms": 1.709,
"diffs": {},
"actual_category": "movie",
"actual_title": "Gladiator",
"actual_destination_subpath": "Movies/Gladiator (2000)/Gladiator (2000) [Remastered].mkv",
"exception": null
},
{
"id": "MOVIE-13",
"domain": "Movies",
"filename": "Seven.Samurai.1954.Criterion.Collection.1080p.BluRay.mkv",
"edge_case_type": "edition_criterion",
"passed": true,
"duration_ms": 1.723,
"diffs": {},
"actual_category": "movie",
"actual_title": "Seven Samurai",
"actual_destination_subpath": "Movies/Seven Samurai (1954)/Seven Samurai (1954) [Criterion].mkv",
"exception": null
},
{
"id": "MOVIE-14",
"domain": "Movies",
"filename": "Blade.Runner.1982.Final.Cut.2160p.UHD.mkv",
"edge_case_type": "edition_final_cut",
"passed": true,
"duration_ms": 1.767,
"diffs": {},
"actual_category": "movie",
"actual_title": "Blade Runner",
"actual_destination_subpath": "Movies/Blade Runner (1982)/Blade Runner (1982) [Final Cut].mkv",
"exception": null
},
{
"id": "SPECIAL-01",
"domain": "Specials & Extras",
"filename": "The.Office.S00E01.The.Outtakes.mkv",
"edge_case_type": "season_00_special",
"passed": true,
"duration_ms": 3.02,
"diffs": {},
"actual_category": "tv",
"actual_title": "The Office",
"actual_destination_subpath": "TV Shows/The Office/Season 00/The Office - S00E01.mkv",
"exception": null
},
{
"id": "SPECIAL-02",
"domain": "Specials & Extras",
"filename": "Doctor.Who.S00E25.The.Day.of.the.Doctor.1080p.mkv",
"edge_case_type": "season_00_high_episode",
"passed": true,
"duration_ms": 2.101,
"diffs": {},
"actual_category": "tv",
"actual_title": "Doctor Who",
"actual_destination_subpath": "TV Shows/Doctor Who/Season 00/Doctor Who - S00E25.mkv",
"exception": null
},
{
"id": "SPECIAL-03",
"domain": "Specials & Extras",
"filename": "Breaking.Bad.S05E00.Special.mkv",
"edge_case_type": "mid_season_special_e00",
"passed": true,
"duration_ms": 1.927,
"diffs": {},
"actual_category": "tv",
"actual_title": "Breaking Bad",
"actual_destination_subpath": "TV Shows/Breaking Bad/Season 05/Breaking Bad - S05E00.mkv",
"exception": null
},
{
"id": "SPECIAL-04",
"domain": "Specials & Extras",
"filename": "Inception.2010-behindthescenes.mkv",
"edge_case_type": "movie_extra_behind_the_scenes",
"passed": true,
"duration_ms": 1.712,
"diffs": {},
"actual_category": "movie",
"actual_title": "Inception",
"actual_destination_subpath": "Movies/Inception (2010)/Inception (2010)-behindthescenes.mkv",
"exception": null
},
{
"id": "SPECIAL-05",
"domain": "Specials & Extras",
"filename": "The.Matrix.1999-featurette.mkv",
"edge_case_type": "movie_extra_featurette",
"passed": true,
"duration_ms": 1.697,
"diffs": {},
"actual_category": "movie",
"actual_title": "The Matrix",
"actual_destination_subpath": "Movies/The Matrix (1999)/The Matrix (1999)-featurette.mkv",
"exception": null
},
{
"id": "SPECIAL-06",
"domain": "Specials & Extras",
"filename": "Interstellar.2014-deleted.mkv",
"edge_case_type": "movie_extra_deleted_scenes",
"passed": true,
"duration_ms": 1.684,
"diffs": {},
"actual_category": "movie",
"actual_title": "Interstellar",
"actual_destination_subpath": "Movies/Interstellar (2014)/Interstellar (2014)-deleted.mkv",
"exception": null
},
{
"id": "SPECIAL-07",
"domain": "Specials & Extras",
"filename": "Interstellar.2014-trailer.mp4",
"edge_case_type": "movie_extra_trailer",
"passed": true,
"duration_ms": 1.684,
"diffs": {},
"actual_category": "movie",
"actual_title": "Interstellar",
"actual_destination_subpath": "Movies/Interstellar (2014)/Interstellar (2014)-trailer.mp4",
"exception": null
},
{
"id": "SPECIAL-08",
"domain": "Specials & Extras",
"filename": "Game.of.Thrones.S00E02.A.Day.in.the.Life.mkv",
"edge_case_type": "tv_special_episode_title",
"passed": true,
"duration_ms": 1.701,
"diffs": {},
"actual_category": "tv",
"actual_title": "Game of Thrones",
"actual_destination_subpath": "TV Shows/Game of Thrones/Season 00/Game of Thrones - S00E02.mkv",
"exception": null
},
{
"id": "SPECIAL-09",
"domain": "Specials & Extras",
"filename": "Sherlock.S00E01.Many.Happy.Returns.mkv",
"edge_case_type": "tv_special_prequel",
"passed": true,
"duration_ms": 1.708,
"diffs": {},
"actual_category": "tv",
"actual_title": "Sherlock",
"actual_destination_subpath": "TV Shows/Sherlock/Season 00/Sherlock - S00E01.mkv",
"exception": null
},
{
"id": "DAILY-01",
"domain": "Daily / Dated Shows",
"filename": "The.Daily.Show.2024-01-15.1080p.HDTV.mkv",
"edge_case_type": "iso_dated_tv",
"passed": true,
"duration_ms": 1.711,
"diffs": {},
"actual_category": "tv",
"actual_title": "The Daily Show",
"actual_destination_subpath": "TV Shows/The Daily Show/Season 2024/The Daily Show - 2024-01-15.mkv",
"exception": null
},
{
"id": "DAILY-02",
"domain": "Daily / Dated Shows",
"filename": "The.Tonight.Show.Starring.Jimmy.Fallon.2024.03.12.720p.mkv",
"edge_case_type": "dot_dated_tv",
"passed": true,
"duration_ms": 1.761,
"diffs": {},
"actual_category": "tv",
"actual_title": "The Tonight Show Starring Jimmy Fallon",
"actual_destination_subpath": "TV Shows/The Tonight Show Starring Jimmy Fallon/Season 2024/The Tonight Show Starring Jimmy Fallon - 2024-03-12.mkv",
"exception": null
},
{
"id": "DAILY-03",
"domain": "Daily / Dated Shows",
"filename": "Last.Week.Tonight.with.John.Oliver.2023-11-05.1080p.mkv",
"edge_case_type": "weekly_dated_show",
"passed": true,
"duration_ms": 1.726,
"diffs": {},
"actual_category": "tv",
"actual_title": "Last Week Tonight with John Oliver",
"actual_destination_subpath": "TV Shows/Last Week Tonight with John Oliver/Season 2023/Last Week Tonight with John Oliver - 2023-11-05.mkv",
"exception": null
},
{
"id": "DAILY-04",
"domain": "Daily / Dated Shows",
"filename": "Late.Night.with.Seth.Meyers.2024_02_20.720p.mkv",
"edge_case_type": "underscore_dated_tv",
"passed": true,
"duration_ms": 1.725,
"diffs": {},
"actual_category": "tv",
"actual_title": "Late Night with Seth Meyers",
"actual_destination_subpath": "TV Shows/Late Night with Seth Meyers/Season 2024/Late Night with Seth Meyers - 2024-02-20.mkv",
"exception": null
},
{
"id": "DAILY-05",
"domain": "Daily / Dated Shows",
"filename": "The Daily - 2026-03-12 - The Sunday Read.mp3",
"edge_case_type": "dated_podcast",
"passed": true,
"duration_ms": 1.711,
"diffs": {},
"actual_category": "podcast",
"actual_title": "The Sunday Read",
"actual_destination_subpath": "Podcasts/The Daily/2026/The Daily - 2026-03-12 - The Sunday Read.mp3",
"exception": null
},
{
"id": "DAILY-06",
"domain": "Daily / Dated Shows",
"filename": "Jimmy.Kimmel.Live.2024-04-18.720p.HDTV.mkv",
"edge_case_type": "daily_show_iso",
"passed": true,
"duration_ms": 1.713,
"diffs": {},
"actual_category": "tv",
"actual_title": "Jimmy Kimmel Live",
"actual_destination_subpath": "TV Shows/Jimmy Kimmel Live/Season 2024/Jimmy Kimmel Live - 2024-04-18.mkv",
"exception": null
},
{
"id": "DAILY-07",
"domain": "Daily / Dated Shows",
"filename": "PBS.NewsHour.2024.05.01.720p.mkv",
"edge_case_type": "news_broadcast_dot_date",
"passed": true,
"duration_ms": 1.709,
"diffs": {},
"actual_category": "tv",
"actual_title": "PBS NewsHour",
"actual_destination_subpath": "TV Shows/PBS NewsHour/Season 2024/PBS NewsHour - 2024-05-01.mkv",
"exception": null
},
{
"id": "DAILY-08",
"domain": "Daily / Dated Shows",
"filename": "The.Late.Show.with.Stephen.Colbert.2024-02-14.1080p.mkv",
"edge_case_type": "late_night_show_iso",
"passed": true,
"duration_ms": 1.724,
"diffs": {},
"actual_category": "tv",
"actual_title": "The Late Show with Stephen Colbert",
"actual_destination_subpath": "TV Shows/The Late Show with Stephen Colbert/Season 2024/The Late Show with Stephen Colbert - 2024-02-14.mkv",
"exception": null
},
{
"id": "DAILY-09",
"domain": "Daily / Dated Shows",
"filename": "NPR.News.Now.2024-06-10.mp3",
"edge_case_type": "podcast_daily_news",
"passed": true,
"duration_ms": 1.708,
"diffs": {},
"actual_category": "podcast",
"actual_title": "NPR News Now",
"actual_destination_subpath": "Podcasts/NPR News Now/2024/NPR News Now - 2024-06-10.mp3",
"exception": null
},
{
"id": "MESSY-01",
"domain": "Messy & Complex",
"filename": "Amélie.2001.PROPER.REMASTERED.1080p.BluRay.x264-CiNEFiLE.mkv",
"edge_case_type": "unicode_accented_characters",
"passed": true,
"duration_ms": 1.764,
"diffs": {},
"actual_category": "movie",
"actual_title": "Amélie",
"actual_destination_subpath": "Movies/Amélie (2001)/Amélie (2001) [Remastered].mkv",
"exception": null
},
{
"id": "MESSY-02",
"domain": "Messy & Complex",
"filename": "Wolfs.2024.1080p.Apple.TV.WEB-DL.DDP5.1.Atmos.H.264.mkv",
"edge_case_type": "apple_tv_scene_tag",
"passed": true,
"duration_ms": 1.758,
"diffs": {},
"actual_category": "movie",
"actual_title": "Wolfs",
"actual_destination_subpath": "Movies/Wolfs (2024)/Wolfs (2024).mkv",
"exception": null
},
{
"id": "MESSY-03",
"domain": "Messy & Complex",
"filename": "Gladiator.II.2024.1080p.HDTV.x264-[rartv].mkv",
"edge_case_type": "roman_numeral_title_bracket_group",
"passed": true,
"duration_ms": 1.742,
"diffs": {},
"actual_category": "movie",
"actual_title": "Gladiator II",
"actual_destination_subpath": "Movies/Gladiator II (2024)/Gladiator II (2024).mkv",
"exception": null
},
{
"id": "MESSY-04",
"domain": "Messy & Complex",
"filename": "Interstellar.1920x1080.mkv",
"edge_case_type": "resolution_dimensions",
"passed": true,
"duration_ms": 1.712,
"diffs": {},
"actual_category": "movie",
"actual_title": "Interstellar",
"actual_destination_subpath": "Movies/Interstellar/Interstellar.mkv",
"exception": null
},
{
"id": "MESSY-05",
"domain": "Messy & Complex",
"filename": "[YTS.MX] Movie Title - 2024 [1080p].mkv",
"edge_case_type": "bracket_group_movie_not_anime",
"passed": true,
"duration_ms": 1.747,
"diffs": {},
"actual_category": "movie",
"actual_title": "Movie Title",
"actual_destination_subpath": "Movies/Movie Title (2024)/Movie Title (2024).mkv",
"exception": null
},
{
"id": "MESSY-06",
"domain": "Messy & Complex",
"filename": "The.Dark.Knight.2008.1080p.forced.srt",
"edge_case_type": "subtitle_forced_tag",
"passed": true,
"duration_ms": 2.06,
"diffs": {},
"actual_category": "subtitle",
"actual_title": "The Dark Knight",
"actual_destination_subpath": "Movies/The Dark Knight (2008)/The Dark Knight (2008).forced.srt",
"exception": null
},
{
"id": "MESSY-07",
"domain": "Messy & Complex",
"filename": "Show: \"Special\" <Episode> | 1?.mkv",
"edge_case_type": "illegal_characters",
"passed": true,
"duration_ms": 1.807,
"diffs": {},
"actual_category": "tv",
"actual_title": "Show Special Episode 1",
"actual_destination_subpath": "TV Shows/Show Special Episode 1/Season 01/Show Special Episode 1 - S01E01.mkv",
"exception": null
},
{
"id": "MESSY-08",
"domain": "Messy & Complex",
"filename": "CON.mp4",
"edge_case_type": "windows_reserved_device_name",
"passed": true,
"duration_ms": 1.845,
"diffs": {},
"actual_category": "home_video",
"actual_title": "CON",
"actual_destination_subpath": "Home Videos/2026/2026-01 - Event/_CON.mp4",
"exception": null
},
{
"id": "MESSY-09",
"domain": "Messy & Complex",
"filename": " Messy Show . S01E01 . 1080p .mkv",
"edge_case_type": "irregular_whitespace_and_dots",
"passed": true,
"duration_ms": 1.738,
"diffs": {},
"actual_category": "tv",
"actual_title": "Messy Show",
"actual_destination_subpath": "TV Shows/Messy Show/Season 01/Messy Show - S01E01.mkv",
"exception": null
},
{
"id": "MESSY-10",
"domain": "Messy & Complex",
"filename": "Show_Name__2022__S02E03__HDTV.mkv",
"edge_case_type": "consecutive_underscores",
"passed": true,
"duration_ms": 1.732,
"diffs": {},
"actual_category": "tv",
"actual_title": "Show Name",
"actual_destination_subpath": "TV Shows/Show Name/Season 02/Show Name - S02E03.mkv",
"exception": null
}
]
}
+344
View File
@@ -0,0 +1,344 @@
"""Standalone CLI Benchmark Runner & Diagnostic Reporter (Requirement R2).
Executable via:
.venv/bin/python -m tests.benchmark.runner
.venv/bin/python tests/benchmark/runner.py
Computes per-domain pass/fail statistics, execution timing, and field diffs for
any failed case. Renders a clean terminal summary table and exports
tests/benchmark/benchmark_summary.json.
"""
from __future__ import annotations
import argparse
from datetime import datetime, timezone
import json
from pathlib import Path
import sys
from typing import Any, Dict, List, Optional
# Ensure project root is in sys.path
_PROJECT_ROOT = Path(__file__).resolve().parent.parent.parent
if str(_PROJECT_ROOT) not in sys.path:
sys.path.insert(0, str(_PROJECT_ROOT))
try:
from tests.benchmark.benchmark_cases import (
BENCHMARK_CASES,
BenchmarkCase,
CaseResult,
evaluate_benchmark_case,
)
except ImportError:
from benchmark_cases import ( # type: ignore
BENCHMARK_CASES,
BenchmarkCase,
CaseResult,
evaluate_benchmark_case,
)
DEFAULT_JSON_PATH = Path(__file__).parent / "benchmark_summary.json"
class BenchmarkRunner:
"""Orchestrates execution, diagnostic collection, and reporting for benchmark cases."""
def __init__(
self,
cases: Optional[List[BenchmarkCase]] = None,
json_output_path: Path = DEFAULT_JSON_PATH,
verbose: bool = False,
show_diffs: bool = True,
strict: bool = False,
):
self.cases = cases or BENCHMARK_CASES
self.json_output_path = json_output_path
self.verbose = verbose
self.show_diffs = show_diffs
self.strict = strict
self.results: List[CaseResult] = []
def run(self) -> Dict[str, Any]:
"""Execute all configured benchmark cases and collect results."""
self.results.clear()
for case in self.cases:
result = evaluate_benchmark_case(case)
self.results.append(result)
summary_data = self._build_summary_data()
self._export_json(summary_data)
return summary_data
def _build_summary_data(self) -> Dict[str, Any]:
total_cases = len(self.results)
passed_cases = sum(1 for r in self.results if r.passed)
failed_cases = total_cases - passed_cases
pass_rate = (passed_cases / total_cases * 100.0) if total_cases else 0.0
total_duration_ms = sum(r.duration_ms for r in self.results)
# Domain breakdown
domain_stats: Dict[str, Dict[str, Any]] = {}
for r in self.results:
d = domain_stats.setdefault(
r.domain,
{
"total": 0,
"passed": 0,
"failed": 0,
"pass_rate_pct": 0.0,
"total_duration_ms": 0.0,
"avg_duration_ms": 0.0,
},
)
d["total"] += 1
if r.passed:
d["passed"] += 1
else:
d["failed"] += 1
d["total_duration_ms"] += r.duration_ms
for d in domain_stats.values():
if d["total"] > 0:
d["pass_rate_pct"] = round(d["passed"] / d["total"] * 100.0, 1)
d["avg_duration_ms"] = round(d["total_duration_ms"] / d["total"], 2)
# Failures list
failures = [
{
"id": r.case_id,
"domain": r.domain,
"filename": r.filename,
"edge_case_type": r.edge_case_type,
"diffs": r.diffs,
"actual_category": r.actual_category,
"actual_title": r.actual_title,
"actual_destination_subpath": r.actual_destination_subpath,
"duration_ms": r.duration_ms,
}
for r in self.results
if not r.passed
]
# Serialized results
all_results_data = [
{
"id": r.case_id,
"domain": r.domain,
"filename": r.filename,
"edge_case_type": r.edge_case_type,
"passed": r.passed,
"duration_ms": r.duration_ms,
"diffs": r.diffs,
"actual_category": r.actual_category,
"actual_title": r.actual_title,
"actual_destination_subpath": r.actual_destination_subpath,
"exception": r.exception,
}
for r in self.results
]
return {
"metadata": {
"timestamp": datetime.now(timezone.utc).isoformat(),
"python_version": sys.version.split()[0],
"total_duration_ms": round(total_duration_ms, 2),
"runner": "MediaSorter Offline E2E Benchmark Runner",
},
"summary": {
"total_cases": total_cases,
"passed_cases": passed_cases,
"failed_cases": failed_cases,
"pass_rate_pct": round(pass_rate, 2),
},
"domains": domain_stats,
"failures": failures,
"all_results": all_results_data,
}
def _export_json(self, data: Dict[str, Any]) -> None:
self.json_output_path.parent.mkdir(parents=True, exist_ok=True)
with open(self.json_output_path, "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
def print_terminal_report(self, summary_data: Dict[str, Any]) -> None:
"""Render a styled terminal summary table and failure details."""
try:
self._print_rich_report(summary_data)
except Exception:
self._print_plain_report(summary_data)
def _print_rich_report(self, data: Dict[str, Any]) -> None:
from rich.console import Console
from rich.markup import escape
from rich.panel import Panel
from rich.table import Table
console = Console()
console.print()
console.print(
Panel.fit(
"[bold cyan]Media Sorter E2E Benchmark Test Suite & Offline Runner[/bold cyan]\n"
f"[dim]Timestamp: {data['metadata']['timestamp']} | Total Cases: {data['summary']['total_cases']}[/dim]",
border_style="cyan",
)
)
table = Table(title="Benchmark Results by Domain", show_footer=True)
table.add_column("Domain", style="bold white", footer="Total / Overall")
table.add_column("Total", justify="right", footer=str(data["summary"]["total_cases"]))
table.add_column("Passed", justify="right", style="green", footer=str(data["summary"]["passed_cases"]))
table.add_column("Failed", justify="right", style="red", footer=str(data["summary"]["failed_cases"]))
table.add_column(
"Pass Rate",
justify="right",
style="bold yellow",
footer=f"{data['summary']['pass_rate_pct']:.1f}%",
)
table.add_column(
"Avg Time (ms)",
justify="right",
style="dim",
footer=f"{data['metadata']['total_duration_ms']:.1f}ms total",
)
for domain, stats in data["domains"].items():
rate = stats["pass_rate_pct"]
rate_style = "green" if rate == 100.0 else ("yellow" if rate >= 50.0 else "red")
table.add_row(
domain,
str(stats["total"]),
str(stats["passed"]),
str(stats["failed"]),
f"[{rate_style}]{rate:.1f}%[/{rate_style}]",
f"{stats['avg_duration_ms']:.2f}ms",
)
console.print(table)
console.print()
# Print failures if requested
if self.show_diffs and data["failures"]:
console.print(f"[bold red]Diagnostic Gap Details ({len(data['failures'])} failures pending M2/M3):[/bold red]")
for item in data["failures"]:
esc_filename = escape(item["filename"])
console.print(
f"\n [bold red]✖ [{item['id']}][/bold red] [bold white]{esc_filename}[/bold white] "
f"([dim]{item['domain']} / {item['edge_case_type']}[/dim])"
)
for field_name, diff in item["diffs"].items():
exp_val = escape(repr(diff["expected"]))
act_val = escape(repr(diff["actual"]))
console.print(
f" [yellow]• {field_name}:[/yellow] "
f"expected=[green]{exp_val}[/green], got=[red]{act_val}[/red]"
)
console.print()
console.print(
f"[dim]Summary report exported to:[/dim] [cyan]{self.json_output_path.resolve()}[/cyan]\n"
)
def _print_plain_report(self, data: Dict[str, Any]) -> None:
print("=" * 80)
print(" Media Sorter E2E Benchmark Test Suite & Offline Runner")
print(f" Timestamp: {data['metadata']['timestamp']} | Total: {data['summary']['total_cases']}")
print("=" * 80)
print(f"{'Domain':<25} {'Total':>8} {'Passed':>8} {'Failed':>8} {'Pass %':>10} {'Avg (ms)':>10}")
print("-" * 80)
for domain, stats in data["domains"].items():
print(
f"{domain:<25} {stats['total']:>8} {stats['passed']:>8} {stats['failed']:>8} "
f"{stats['pass_rate_pct']:>9.1f}% {stats['avg_duration_ms']:>10.2f}"
)
print("-" * 80)
summ = data["summary"]
print(
f"{'Total / Overall':<25} {summ['total_cases']:>8} {summ['passed_cases']:>8} {summ['failed_cases']:>8} "
f"{summ['pass_rate_pct']:>9.1f}% {data['metadata']['total_duration_ms']:>10.2f}ms"
)
print("=" * 80)
if self.show_diffs and data["failures"]:
print(f"\nDiagnostic Gap Details ({len(data['failures'])} failures pending M2/M3):")
for item in data["failures"]:
print(f" * [{item['id']}] {item['filename']} ({item['domain']} / {item['edge_case_type']})")
for k, v in item["diffs"].items():
print(f" - {k}: expected={v['expected']!r}, got={v['actual']!r}")
print(f"\nSummary JSON exported to: {self.json_output_path.resolve()}\n")
def main() -> int:
parser = argparse.ArgumentParser(
description="Offline E2E Benchmark Runner for Media Sorter Pattern Recognition."
)
parser.add_argument(
"--domain",
type=str,
default=None,
help="Filter execution by specific domain (e.g. 'Anime', 'Standard TV', 'Movies').",
)
parser.add_argument(
"--case",
type=str,
default=None,
help="Filter execution by specific BenchmarkCase ID (e.g. 'TV-01', 'ANIME-04').",
)
parser.add_argument(
"--json-output",
type=str,
default=str(DEFAULT_JSON_PATH),
help="Output JSON summary destination path.",
)
parser.add_argument(
"--verbose",
"-v",
action="store_true",
help="Enable verbose output reporting.",
)
parser.add_argument(
"--no-diffs",
action="store_true",
help="Suppress detailed failure diff output.",
)
parser.add_argument(
"--strict",
action="store_true",
help="Exit with non-zero status code if any benchmark case fails.",
)
args = parser.parse_args()
selected_cases = BENCHMARK_CASES
if args.domain:
selected_cases = [c for c in selected_cases if c.domain.lower() == args.domain.lower()]
if not selected_cases:
print(f"Error: No benchmark cases match domain '{args.domain}'", file=sys.stderr)
return 2
if args.case:
selected_cases = [c for c in selected_cases if c.id.upper() == args.case.upper()]
if not selected_cases:
print(f"Error: No benchmark case matches ID '{args.case}'", file=sys.stderr)
return 2
runner = BenchmarkRunner(
cases=selected_cases,
json_output_path=Path(args.json_output),
verbose=args.verbose,
show_diffs=not args.no_diffs,
strict=args.strict,
)
summary_data = runner.run()
runner.print_terminal_report(summary_data)
if args.strict and summary_data["summary"]["failed_cases"] > 0:
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
+56
View File
@@ -0,0 +1,56 @@
"""Automated E2E Benchmark Test Suite (Requirement R2).
Executes 100% offline with zero network connectivity and zero mutations to real disk storage.
Parametrized across all 60+ benchmark cases across the 6 media domains:
1. Standard TV
2. Anime
3. Movies
4. Specials & Extras
5. Daily / Dated Shows
6. Messy & Complex
"""
from __future__ import annotations
import sys
from pathlib import Path
import pytest
# Ensure project root is in sys.path regardless of how pytest is invoked
_PROJECT_ROOT = Path(__file__).resolve().parent.parent.parent
if str(_PROJECT_ROOT) not in sys.path:
sys.path.insert(0, str(_PROJECT_ROOT))
try:
from tests.benchmark.benchmark_cases import (
BENCHMARK_CASES,
BenchmarkCase,
evaluate_benchmark_case,
)
except ImportError:
from benchmark_cases import ( # type: ignore
BENCHMARK_CASES,
BenchmarkCase,
evaluate_benchmark_case,
)
@pytest.mark.parametrize("case", BENCHMARK_CASES, ids=lambda c: c.id)
def test_benchmark(case: BenchmarkCase) -> None:
"""Execute end-to-end benchmark test for an individual media file pattern.
Tests FilenameTokenizer tokenization, MediaClassifier classification, and
MediaNamer destination path generation against expected baseline values.
"""
result = evaluate_benchmark_case(case)
if not result.passed:
error_lines = [
f"Benchmark case {case.id} ({case.domain} - {case.edge_case_type}) failed:",
f" Filename: {case.filename}",
]
for field_name, mismatch in result.diffs.items():
error_lines.append(
f" * {field_name}: expected={mismatch['expected']!r}, got={mismatch['actual']!r}"
)
pytest.fail("\n".join(error_lines))
+319
View File
@@ -0,0 +1,319 @@
import builtins
import io
import ipaddress
import os
from pathlib import Path
import shutil
import socket
import pytest
PROTECTED_ROOTS = [
Path("/md0/jdownloads"),
Path("/md0/movies1"),
Path("/md0/tv1"),
]
PROTECTED_MATCHERS = []
for p in PROTECTED_ROOTS:
PROTECTED_MATCHERS.append(p)
try:
PROTECTED_MATCHERS.append(p.resolve())
except Exception:
pass
def check_target(target):
"""Check if target path resolves into or is inside protected production directories."""
if target is None or isinstance(target, int):
return
try:
if isinstance(target, bytes):
t_str = os.fsdecode(target)
else:
t_str = str(target)
p = Path(t_str)
p_resolved = p.resolve()
except Exception:
return
for root in PROTECTED_MATCHERS:
try:
if p_resolved == root or p_resolved.is_relative_to(root):
raise RuntimeError(
f"FILESYSTEM SAFETY TRAP: Forbidden write/delete operation targeting production path '{target}' in test execution!"
)
except RuntimeError:
raise
except Exception:
pass
try:
if p == root or p.is_relative_to(root):
raise RuntimeError(
f"FILESYSTEM SAFETY TRAP: Forbidden write/delete operation targeting production path '{target}' in test execution!"
)
except RuntimeError:
raise
except Exception:
pass
@pytest.fixture(autouse=True, scope="session")
def protect_production_filesystem():
"""Active safety trap that monkeypatches filesystem modification operations in
os, shutil, pathlib.Path, and open (for write/append/create modes).
Inspects all target paths: if any path resolves into or is inside /md0/jdownloads,
/md0/movies1, or /md0/tv1, raises RuntimeError.
"""
mp = pytest.MonkeyPatch()
# 1. Builtins & IO open
orig_builtin_open = builtins.open
def safe_builtin_open(file, *args, **kwargs):
mode = "r"
if args:
mode = args[0]
elif "mode" in kwargs:
mode = kwargs["mode"]
if any(m in mode for m in ("w", "a", "x", "+")):
check_target(file)
return orig_builtin_open(file, *args, **kwargs)
mp.setattr(builtins, "open", safe_builtin_open)
mp.setattr(io, "open", safe_builtin_open)
# 2. pathlib.Path methods
orig_path_open = Path.open
orig_path_write_text = Path.write_text
orig_path_write_bytes = Path.write_bytes
orig_path_unlink = Path.unlink
orig_path_rmdir = Path.rmdir
orig_path_mkdir = Path.mkdir
orig_path_rename = Path.rename
orig_path_replace = Path.replace
orig_path_touch = Path.touch
def safe_path_open(self, *args, **kwargs):
mode = "r"
if args:
mode = args[0]
elif "mode" in kwargs:
mode = kwargs["mode"]
if any(m in mode for m in ("w", "a", "x", "+")):
check_target(self)
return orig_path_open(self, *args, **kwargs)
def safe_path_write_text(self, *args, **kwargs):
check_target(self)
return orig_path_write_text(self, *args, **kwargs)
def safe_path_write_bytes(self, *args, **kwargs):
check_target(self)
return orig_path_write_bytes(self, *args, **kwargs)
def safe_path_unlink(self, *args, **kwargs):
check_target(self)
return orig_path_unlink(self, *args, **kwargs)
def safe_path_rmdir(self, *args, **kwargs):
check_target(self)
return orig_path_rmdir(self, *args, **kwargs)
def safe_path_mkdir(self, *args, **kwargs):
check_target(self)
return orig_path_mkdir(self, *args, **kwargs)
def safe_path_rename(self, target, *args, **kwargs):
check_target(self)
check_target(target)
return orig_path_rename(self, target, *args, **kwargs)
def safe_path_replace(self, target, *args, **kwargs):
check_target(self)
check_target(target)
return orig_path_replace(self, target, *args, **kwargs)
def safe_path_touch(self, *args, **kwargs):
check_target(self)
return orig_path_touch(self, *args, **kwargs)
mp.setattr(Path, "open", safe_path_open)
mp.setattr(Path, "write_text", safe_path_write_text)
mp.setattr(Path, "write_bytes", safe_path_write_bytes)
mp.setattr(Path, "unlink", safe_path_unlink)
mp.setattr(Path, "rmdir", safe_path_rmdir)
mp.setattr(Path, "mkdir", safe_path_mkdir)
mp.setattr(Path, "rename", safe_path_rename)
mp.setattr(Path, "replace", safe_path_replace)
mp.setattr(Path, "touch", safe_path_touch)
# 3. os functions
orig_os_remove = os.remove
orig_os_unlink = os.unlink
orig_os_rmdir = os.rmdir
orig_os_mkdir = os.mkdir
orig_os_makedirs = os.makedirs
orig_os_rename = os.rename
orig_os_replace = os.replace
orig_os_open = os.open
orig_os_truncate = os.truncate
def safe_os_remove(path, *args, **kwargs):
check_target(path)
return orig_os_remove(path, *args, **kwargs)
def safe_os_unlink(path, *args, **kwargs):
check_target(path)
return orig_os_unlink(path, *args, **kwargs)
def safe_os_rmdir(path, *args, **kwargs):
check_target(path)
return orig_os_rmdir(path, *args, **kwargs)
def safe_os_mkdir(path, *args, **kwargs):
check_target(path)
return orig_os_mkdir(path, *args, **kwargs)
def safe_os_makedirs(name, *args, **kwargs):
check_target(name)
return orig_os_makedirs(name, *args, **kwargs)
def safe_os_rename(src, dst, *args, **kwargs):
check_target(src)
check_target(dst)
return orig_os_rename(src, dst, *args, **kwargs)
def safe_os_replace(src, dst, *args, **kwargs):
check_target(src)
check_target(dst)
return orig_os_replace(src, dst, *args, **kwargs)
def safe_os_open(path, flags, *args, **kwargs):
if flags & (os.O_WRONLY | os.O_RDWR | os.O_CREAT | os.O_TRUNC | os.O_APPEND):
check_target(path)
return orig_os_open(path, flags, *args, **kwargs)
def safe_os_truncate(path, *args, **kwargs):
check_target(path)
return orig_os_truncate(path, *args, **kwargs)
mp.setattr(os, "remove", safe_os_remove)
mp.setattr(os, "unlink", safe_os_unlink)
mp.setattr(os, "rmdir", safe_os_rmdir)
mp.setattr(os, "mkdir", safe_os_mkdir)
mp.setattr(os, "makedirs", safe_os_makedirs)
mp.setattr(os, "rename", safe_os_rename)
mp.setattr(os, "replace", safe_os_replace)
mp.setattr(os, "open", safe_os_open)
mp.setattr(os, "truncate", safe_os_truncate)
# 4. shutil functions
orig_shutil_rmtree = shutil.rmtree
orig_shutil_move = shutil.move
orig_shutil_copy = shutil.copy
orig_shutil_copy2 = shutil.copy2
orig_shutil_copyfile = shutil.copyfile
orig_shutil_copytree = shutil.copytree
def safe_shutil_rmtree(path, *args, **kwargs):
check_target(path)
return orig_shutil_rmtree(path, *args, **kwargs)
def safe_shutil_move(src, dst, *args, **kwargs):
check_target(src)
check_target(dst)
return orig_shutil_move(src, dst, *args, **kwargs)
def safe_shutil_copy(src, dst, *args, **kwargs):
check_target(dst)
return orig_shutil_copy(src, dst, *args, **kwargs)
def safe_shutil_copy2(src, dst, *args, **kwargs):
check_target(dst)
return orig_shutil_copy2(src, dst, *args, **kwargs)
def safe_shutil_copyfile(src, dst, *args, **kwargs):
check_target(dst)
return orig_shutil_copyfile(src, dst, *args, **kwargs)
def safe_shutil_copytree(src, dst, *args, **kwargs):
check_target(dst)
return orig_shutil_copytree(src, dst, *args, **kwargs)
mp.setattr(shutil, "rmtree", safe_shutil_rmtree)
mp.setattr(shutil, "move", safe_shutil_move)
mp.setattr(shutil, "copy", safe_shutil_copy)
mp.setattr(shutil, "copy2", safe_shutil_copy2)
mp.setattr(shutil, "copyfile", safe_shutil_copyfile)
mp.setattr(shutil, "copytree", safe_shutil_copytree)
yield
mp.undo()
@pytest.fixture(autouse=True, scope="function")
def isolate_test_environment():
"""Saves os.environ before test and restores it after.
Strips existing CONFIDENCE_THRESHOLD, DOWNLOADS_DIR, MOVIES_DIR,
SHOWS_DIR, ANIME_DIR, SOURCE_DIR, TV_DIR, DRY_RUN, ACTION, and MEDIA_SORTER_*
variables so Settings() initializes with clean test defaults and prevents
cross-test contamination.
"""
saved_env = dict(os.environ)
keys_to_strip = [
"CONFIDENCE_THRESHOLD",
"DOWNLOADS_DIR",
"MOVIES_DIR",
"SHOWS_DIR",
"ANIME_DIR",
"SOURCE_DIR",
"TV_DIR",
"DRY_RUN",
"ACTION",
]
for key in keys_to_strip:
os.environ.pop(key, None)
for key in list(os.environ.keys()):
if key.startswith("MEDIA_SORTER_"):
os.environ.pop(key, None)
try:
yield
finally:
os.environ.clear()
os.environ.update(saved_env)
@pytest.fixture(autouse=True, scope="session")
def block_external_network():
"""Monkeypatches socket.socket.connect to prevent outbound internet network
requests during tests (allow loopback/localhost and unix domain sockets).
"""
mp = pytest.MonkeyPatch()
orig_connect = socket.socket.connect
def safe_connect(self, address):
# Allow Unix domain sockets
if hasattr(socket, "AF_UNIX") and self.family == socket.AF_UNIX:
return orig_connect(self, address)
if isinstance(address, (str, bytes)):
return orig_connect(self, address)
# For INET/INET6 sockets, address is (host, port, ...)
if isinstance(address, tuple) and len(address) >= 1:
host = address[0]
if host in ("localhost", "127.0.0.1", "::1", "0.0.0.0"):
return orig_connect(self, address)
try:
ip = ipaddress.ip_address(host)
if ip.is_loopback:
return orig_connect(self, address)
except ValueError:
pass
raise RuntimeError(
f"NETWORK ACCESS TRAP: Outbound network connection to {address} blocked during tests!"
)
mp.setattr(socket.socket, "connect", safe_connect)
yield
mp.undo()
+102
View File
@@ -0,0 +1,102 @@
import os
import time
from pathlib import Path
import pytest
from media_sorter.config import Settings
from media_sorter.db import get_db_session, init_db
from media_sorter.models import BatchRecord, Operation, QuarantineRecord
from media_sorter.sorter import MediaSorterApp
@pytest.fixture
def library_environment(tmp_path: Path):
incoming = tmp_path / "incoming"
organized = tmp_path / "organized"
db_file = tmp_path / "media_sorter.db"
incoming.mkdir()
organized.mkdir()
settings = Settings()
settings.storage.source_dirs = [str(incoming)]
settings.storage.destination_base = str(organized)
settings.storage.destination_dirs.movies = "Movies"
settings.storage.destination_dirs.tv = "TV Shows"
settings.database.path = str(db_file)
settings.general.min_file_age_seconds = 0 # Process immediately in test
engine = init_db(db_path=db_file)
return settings, engine, incoming, organized
def test_end_to_end_pipeline(library_environment):
settings, engine, incoming, organized = library_environment
# 1. Populate mock media files in incoming directory
movie_file = incoming / "The.Dark.Knight.2008.1080p.BluRay.x264.mkv"
movie_file.write_bytes(b"\x1aE\xdf\xa3" + b"\x00" * 200) # EBML header
sub_file = incoming / "The.Dark.Knight.2008.1080p.BluRay.x264.en.srt"
sub_file.write_text("1\n00:00:01,000 --> 00:00:03,000\nBatman begins.\n")
tv_file = incoming / "The.Wire.S01E01.Target.720p.mkv"
tv_file.write_bytes(b"\x1aE\xdf\xa3" + b"\x00" * 200)
photo_file = incoming / "IMG_20250620_153022.jpg"
photo_file.write_bytes(b"\xff\xd8\xff\xe0" + b"\x00" * 100) # JPEG header
junk_file = incoming / "unrecognized_sample.xyz"
junk_file.write_bytes(b"\x00\x01\x02\x03\x04")
sorter = MediaSorterApp(settings, engine)
# 2. First Run: Dry-Run
dry_report = sorter.run(dry_run=True)
assert dry_report.dry_run is True
assert dry_report.total_files >= 4
# Ensure source files were not modified
assert movie_file.exists()
assert sub_file.exists()
assert tv_file.exists()
assert photo_file.exists()
assert junk_file.exists()
# 3. Second Run: Live Organization
live_report = sorter.run(dry_run=False)
assert live_report.dry_run is False
assert live_report.moved_files >= 3
# Verify Movie and matched subtitle sidecar
movie_dest_dir = organized / "Movies/The Dark Knight (2008)"
assert movie_dest_dir.exists()
dest_movie = list(movie_dest_dir.glob("*.mkv"))[0]
assert "The Dark Knight" in dest_movie.name
# Subtitle should be alongside movie with .en.srt
dest_sub = list(movie_dest_dir.glob("*.srt"))[0]
assert dest_sub.name.endswith(".en.srt")
# Verify TV show organization
tv_season_dir = organized / "TV Shows/The Wire/Season 01"
assert tv_season_dir.exists()
dest_tv = list(tv_season_dir.glob("*.mkv"))[0]
assert "S01E01" in dest_tv.name
# Verify Photo organization
photo_dest_dir = organized / "Photos/2025/2025-06"
assert photo_dest_dir.exists()
# Verify Quarantine of unknown/low confidence file (flagged in place, not moved)
assert junk_file.exists()
from media_sorter.quarantine import QuarantineManager
with get_db_session(engine) as s:
qm = QuarantineManager(s)
pending = qm.list_pending()
assert any("unrecognized_sample.xyz" in q.src for q in pending)
# 4. Third Step: Rollback
reverted = sorter.rollback(live_report.batch_id)
assert reverted >= 3
# Verify files restored to incoming!
assert movie_file.exists()
assert sub_file.exists()
assert tv_file.exists()
@@ -0,0 +1,76 @@
from pathlib import Path
import pytest
from hypothesis import given, settings, strategies as st
from media_sorter.namer import FORBIDDEN_CHARS_PATTERN, RESERVED_NAMES, sanitize_filename_component
from media_sorter.tokenizer import FilenameTokenizer
@pytest.fixture
def tokenizer():
return FilenameTokenizer()
# -----------------------------------------------------------------------------
# Real-World Messy Scene & International Filenames
# -----------------------------------------------------------------------------
MESSY_CASES = [
(
"[HorribleSubs] Shingeki no Kyojin - 59 [1080p].mkv",
{"title": "Shingeki no Kyojin", "episode": 59, "is_anime": True},
),
(
"[SubsPlease] 葬送のフリーレン - 12 (1080p) [98E7B1A2].mkv",
{"title": "葬送のフリーレン", "episode": 12, "is_anime": True},
),
(
"Amélie.2001.PROPER.REMASTERED.1080p.BluRay.x264-CiNEFiLE.mkv",
{"title": "Amélie", "year": 2001, "resolution": "1080p"},
),
(
"Doctor.Who.2005.S01E01.Rose.720p.HDTV.x264-FoV.mkv",
{"title": "Doctor Who", "year": 2005, "season": 1, "episode": 1},
),
(
"Game of Thrones - 1x09 - Baelor [720p HDTV].mkv",
{"title": "Game of Thrones", "season": 1, "episode": 9},
),
(
"Mission.Impossible.Dead.Reckoning.Part.One.2023.2160p.WEB-DL.DDP5.1.Atmos.DV.HDR.H.265-FLUX.mkv",
{"year": 2023, "resolution": "2160p", "video_codec": "h265"},
),
]
@pytest.mark.parametrize("filename,expected", MESSY_CASES)
def test_real_world_messy_filenames(tokenizer, filename, expected):
tokens = tokenizer.tokenize(Path(filename))
for key, val in expected.items():
assert getattr(tokens, key) == val, f"Failed match for {key} in {filename}"
# -----------------------------------------------------------------------------
# Property-Based Fuzz Testing with Hypothesis
# -----------------------------------------------------------------------------
@given(st.text(min_size=1, max_size=500))
@settings(max_examples=150)
def test_sanitize_filename_component_fuzz(input_text):
result = sanitize_filename_component(input_text, max_length=120)
# 1. Result must never be empty
assert len(result) > 0
# 2. Result must contain no forbidden characters
assert not FORBIDDEN_CHARS_PATTERN.search(result)
# 3. Result must not have leading or trailing dots, spaces, or hyphens
assert not result.startswith((" ", ".", "-"))
assert not result.endswith((" ", ".", "-"))
# 4. Result base name must not be a Windows reserved device name
upper_base = result.split(".")[0].upper()
assert upper_base not in RESERVED_NAMES
# 5. Encoded byte length must stay within requested boundary (plus extension tolerance)
assert len(result.encode("utf-8")) <= 140
+116
View File
@@ -0,0 +1,116 @@
import struct
from pathlib import Path
import pytest
from media_sorter.analyzer import MediaAnalyzer
@pytest.fixture
def analyzer():
return MediaAnalyzer()
def test_flac_metadata_extraction(tmp_path: Path, analyzer):
flac_file = tmp_path / "test.flac"
# Create minimal synthetic FLAC file
header = bytearray(b"fLaC")
# STREAMINFO block (type 0, length 34, not last: block_hdr = 0x00 0x00 0x00 0x22)
streaminfo_data = bytearray(34)
# sample_rate = 44100, channels = 2, total_samples = 44100 * 10
# byte 10..12: sample rate
streaminfo_data[10] = (44100 >> 12) & 0xFF
streaminfo_data[11] = (44100 >> 4) & 0xFF
streaminfo_data[12] = ((44100 & 0x0F) << 4) | (1 << 1) # 2 channels (bits 1..3 = 1)
# total samples = 441000
total_samples = 441000
streaminfo_data[13] = (total_samples >> 32) & 0x0F
streaminfo_data[14] = (total_samples >> 24) & 0xFF
streaminfo_data[15] = (total_samples >> 16) & 0xFF
streaminfo_data[16] = (total_samples >> 8) & 0xFF
streaminfo_data[17] = total_samples & 0xFF
header.extend(b"\x80\x00\x00\x22") # is_last = 1, type = 0, len = 34
header.extend(streaminfo_data)
flac_file.write_bytes(header)
meta = analyzer.analyze(flac_file)
assert meta.container == "flac"
assert meta.has_audio is True
assert meta.duration_seconds == 10.0
def test_mp3_id3v2_extraction(tmp_path: Path, analyzer):
mp3_file = tmp_path / "test.mp3"
header = bytearray(b"ID3\x03\x00\x00") # ID3v2.3
# Build TIT2 frame (Title: Bohemian Rhapsody)
tit2_val = b"\x00Bohemian Rhapsody"
tit2_frame = b"TIT2" + struct.pack(">I", len(tit2_val)) + b"\x00\x00" + tit2_val
# Build TPE1 frame (Artist: Queen)
tpe1_val = b"\x00Queen"
tpe1_frame = b"TPE1" + struct.pack(">I", len(tpe1_val)) + b"\x00\x00" + tpe1_val
tag_content = tit2_frame + tpe1_frame
tag_len = len(tag_content)
# Syncsafe integer for tag size
b0 = (tag_len >> 21) & 0x7F
b1 = (tag_len >> 14) & 0x7F
b2 = (tag_len >> 7) & 0x7F
b3 = tag_len & 0x7F
header.extend(bytes([b0, b1, b2, b3]))
header.extend(tag_content)
# Add dummy MP3 frame sync bytes
header.extend(b"\xff\xfb\x90\x00")
mp3_file.write_bytes(header)
meta = analyzer.analyze(mp3_file)
assert meta.container == "mp3"
assert meta.has_audio is True
assert meta.tags.get("title") == "Bohemian Rhapsody"
assert meta.tags.get("artist") == "Queen"
def test_jpeg_exif_extraction(tmp_path: Path, analyzer):
jpg_file = tmp_path / "test.jpg"
# SOI marker + APP1 marker with Exif
soi = b"\xff\xd8"
exif_header = b"Exif\x00\x00"
tiff_header = b"II\x2a\x00\x08\x00\x00\x00" # Little endian, IFD at 8
# 1 entry in IFD0: DateTimeOriginal (tag 0x9003)
num_entries = struct.pack("<H", 1)
tag_id = struct.pack("<H", 0x9003)
type_ascii = struct.pack("<H", 2)
count = struct.pack("<I", 20)
val_offset = struct.pack("<I", 22) # offset from tiff_header
dt_str = b"2025:06:15 10:30:00\x00"
tiff_body = tiff_header + num_entries + tag_id + type_ascii + count + val_offset + dt_str
app1_len = struct.pack(">H", len(exif_header) + len(tiff_body) + 2)
app1 = b"\xff\xe1" + app1_len + exif_header + tiff_body
jpg_file.write_bytes(soi + app1)
meta = analyzer.analyze(jpg_file)
assert meta.container == "jpeg"
assert meta.tags.get("datetime_original") == "2025:06:15 10:30:00"
def test_png_dimensions(tmp_path: Path, analyzer):
png_file = tmp_path / "test.png"
sig = b"\x89PNG\r\n\x1a\n"
# IHDR chunk: 13 bytes data (width=1920, height=1080)
ihdr_data = struct.pack(">IIBBBBB", 1920, 1080, 8, 2, 0, 0, 0)
ihdr = struct.pack(">I", 13) + b"IHDR" + ihdr_data + b"\x00\x00\x00\x00"
png_file.write_bytes(sig + ihdr)
meta = analyzer.analyze(png_file)
assert meta.container == "png"
assert meta.width == 1920
assert meta.height == 1080
assert meta.resolution_label == "1080p"
def test_archive_detection(tmp_path: Path, analyzer):
zip_file = tmp_path / "test.zip"
zip_file.write_bytes(b"PK\x03\x04\x14\x00\x00\x00")
meta = analyzer.analyze(zip_file)
assert meta.container == "zip"
assert meta.mime_type == "application/zip"
+312
View File
@@ -0,0 +1,312 @@
from pathlib import Path
import pytest
from media_sorter.analyzer import MediaMetadata, StreamInfo
from media_sorter.classifier import MediaClassifier
from media_sorter.providers import MockMetadataProvider, ProviderResult
from media_sorter.scanner import ScannedFile
from media_sorter.tokenizer import FilenameTokenizer, TokenizedFilename
@pytest.fixture
def classifier():
return MediaClassifier(confidence_threshold=0.75)
def test_classify_tv_show(classifier):
scanned = ScannedFile(path=Path("/downloads/Game.of.Thrones.S01E01.1080p.mkv"), size=1000000, mtime=1000.0)
tokens = TokenizedFilename(
raw_name="Game.of.Thrones.S01E01.1080p.mkv",
title="Game of Thrones",
season=1,
episode=1,
is_episodic=True,
)
meta = MediaMetadata(
path=scanned.path,
mime_type="video/x-matroska",
container="mkv",
duration_seconds=3600,
has_video=True,
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "tv"
assert res.confidence >= 0.75
assert res.needs_quarantine is False
def test_classify_anime(classifier):
scanned = ScannedFile(path=Path("/downloads/[SubsPlease] Jujutsu Kaisen - 01 [1080p].mkv"), size=1000000, mtime=1000.0)
tokens = TokenizedFilename(
raw_name="[SubsPlease] Jujutsu Kaisen - 01 [1080p].mkv",
title="Jujutsu Kaisen",
episode=1,
season=1,
group="SubsPlease",
is_anime=True,
is_episodic=True,
)
meta = MediaMetadata(
path=scanned.path,
mime_type="video/x-matroska",
container="mkv",
duration_seconds=1400,
has_video=True,
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "anime"
assert res.confidence >= 0.75
assert res.needs_quarantine is False
def test_classify_movie(classifier):
scanned = ScannedFile(path=Path("/downloads/Interstellar.2014.1080p.mkv"), size=5000000, mtime=1000.0)
tokens = TokenizedFilename(
raw_name="Interstellar.2014.1080p.mkv",
title="Interstellar",
year=2014,
resolution="1080p",
)
meta = MediaMetadata(
path=scanned.path,
mime_type="video/x-matroska",
container="mkv",
duration_seconds=10140, # ~2.8 hours
has_video=True,
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "movie"
assert res.confidence >= 0.75
assert res.needs_quarantine is False
def test_classify_music(classifier):
scanned = ScannedFile(path=Path("/music/01 - Come Together.flac"), size=30000000, mtime=1000.0)
tokens = TokenizedFilename(
raw_name="01 - Come Together.flac",
title="Come Together",
track=1,
is_music=True,
)
meta = MediaMetadata(
path=scanned.path,
mime_type="audio/flac",
container="flac",
duration_seconds=259,
has_audio=True,
has_video=False,
tags={"artist": "The Beatles", "album": "Abbey Road"},
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "music"
assert res.confidence >= 0.75
assert res.needs_quarantine is False
def test_classify_audiobook(classifier):
scanned = ScannedFile(path=Path("/audiobooks/Dune - Part 01.m4b"), size=50000000, mtime=1000.0)
tokens = TokenizedFilename(raw_name="Dune - Part 01.m4b", title="Dune")
meta = MediaMetadata(
path=scanned.path,
mime_type="audio/mp4",
container="m4b",
duration_seconds=28800, # 8 hours
has_audio=True,
has_video=False,
tags={"narrator": "George Guidall"},
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "audiobook"
assert res.confidence >= 0.75
assert res.needs_quarantine is False
def test_low_confidence_triggers_quarantine(classifier):
# Ambiguous video clip with no year, no episode, no metadata
scanned = ScannedFile(path=Path("/incoming/unknown_recording_xyz.mkv"), size=10000, mtime=1000.0)
tokens = TokenizedFilename(raw_name="unknown_recording_xyz.mkv", title="unknown recording xyz")
meta = MediaMetadata(
path=scanned.path,
mime_type="video/x-matroska",
container="mkv",
duration_seconds=120,
has_video=True,
)
res = classifier.classify(scanned, tokens, meta)
assert res.confidence < 0.75
assert res.needs_quarantine is True
assert res.quarantine_reason is not None
def test_unsupported_format_triggers_quarantine(classifier):
scanned = ScannedFile(path=Path("/incoming/corrupt_data.bin"), size=1000, mtime=1000.0)
tokens = TokenizedFilename(raw_name="corrupt_data.bin")
meta = MediaMetadata(path=scanned.path, mime_type="application/octet-stream", container="bin")
res = classifier.classify(scanned, tokens, meta)
assert res.category == "unknown"
assert res.needs_quarantine is True
def test_classify_with_metadata_provider():
mock_prov = MockMetadataProvider(
mock_data={
"movie:oppenheimer": ProviderResult(
canonical_title="Oppenheimer",
year=2023,
media_type="movie",
confidence_boost=0.15,
),
"tv:the last of us": ProviderResult(
canonical_title="The Last of Us",
year=2023,
media_type="tv",
season=1,
episode=3,
episode_title="Long, Long Time",
confidence_boost=0.20,
),
}
)
prov_classifier = MediaClassifier(confidence_threshold=0.75, provider=mock_prov)
# 1. Movie verified with provider
scanned_m = ScannedFile(path=Path("/downloads/Oppenheimer.mkv"), size=1000000, mtime=1000.0)
tokens_m = TokenizedFilename(raw_name="Oppenheimer.mkv", title="Oppenheimer")
meta_m = MediaMetadata(path=scanned_m.path, mime_type="video/x-matroska", container="mkv", duration_seconds=10800, has_video=True)
res_m = prov_classifier.classify(scanned_m, tokens_m, meta_m)
assert res_m.category == "movie"
assert res_m.provider_result is not None
assert res_m.provider_result.year == 2023
# 2. TV Show verified with provider
scanned_tv = ScannedFile(path=Path("/downloads/The.Last.of.Us.S01E03.mkv"), size=1000000, mtime=1000.0)
tokens_tv = TokenizedFilename(raw_name="The.Last.of.Us.S01E03.mkv", title="The Last of Us", season=1, episode=3, is_episodic=True)
meta_tv = MediaMetadata(path=scanned_tv.path, mime_type="video/x-matroska", container="mkv", duration_seconds=4500, has_video=True)
res_tv = prov_classifier.classify(scanned_tv, tokens_tv, meta_tv)
assert res_tv.category == "tv"
assert res_tv.provider_result is not None
assert res_tv.provider_result.episode_title == "Long, Long Time"
def test_classify_podcast(classifier):
scanned = ScannedFile(path=Path("/podcasts/Hardcore History 2023-05-12 Episode 68.mp3"), size=50000000, mtime=1000.0)
tokens = TokenizedFilename(raw_name="Hardcore History 2023-05-12 Episode 68.mp3", title="Episode 68", date_stamp="2023-05-12")
meta = MediaMetadata(
path=scanned.path,
mime_type="audio/mpeg",
container="mp3",
duration_seconds=14400,
has_audio=True,
has_video=False,
tags={"podcast": "Dan Carlin's Hardcore History"},
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "podcast"
assert res.confidence >= 0.75
def test_classify_photo_and_home_video(classifier):
# Photo test
scanned_p = ScannedFile(path=Path("/photos/IMG_20250615_123456.jpg"), size=4000000, mtime=1000.0)
tokens_p = TokenizedFilename(raw_name="IMG_20250615_123456.jpg", is_photo_or_home_video=True, date_stamp="2025-06-15")
meta_p = MediaMetadata(path=scanned_p.path, mime_type="image/jpeg", container="jpeg", tags={"camera_model": "Pixel 9 Pro"})
res_p = classifier.classify(scanned_p, tokens_p, meta_p)
assert res_p.category == "photo"
assert res_p.confidence >= 0.90
# Home Video test
scanned_v = ScannedFile(path=Path("/home_videos/VID_20250615_140000.mp4"), size=20000000, mtime=1000.0)
tokens_v = TokenizedFilename(raw_name="VID_20250615_140000.mp4", is_photo_or_home_video=True, date_stamp="2025-06-15")
meta_v = MediaMetadata(path=scanned_v.path, mime_type="video/mp4", container="mp4", duration_seconds=120, has_video=True)
res_v = classifier.classify(scanned_v, tokens_v, meta_v)
assert res_v.category == "home_video"
def test_classify_archive(classifier):
scanned = ScannedFile(path=Path("/downloads/Season1_Extras.zip"), size=500000000, mtime=1000.0)
tokens = TokenizedFilename(raw_name="Season1_Extras.zip")
meta = MediaMetadata(path=scanned.path, mime_type="application/zip", container="zip")
res = classifier.classify(scanned, tokens, meta)
assert res.category == "archive"
assert res.confidence >= 0.90
def test_classify_movie_with_hdtv_and_rartv(classifier):
scanned = ScannedFile(path=Path("/downloads/Gladiator.II.2024.1080p.HDTV.x264-[rartv].mkv"), size=4000000000, mtime=1000.0)
tokens = TokenizedFilename(
raw_name="Gladiator.II.2024.1080p.HDTV.x264-[rartv].mkv",
title="Gladiator II",
year=2024,
resolution="1080p",
video_codec="x264",
source="HDTV",
group="rartv",
)
meta = MediaMetadata(
path=scanned.path,
mime_type="video/x-matroska",
container="mkv",
duration_seconds=5000, # ~83 minutes
has_video=True,
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "movie"
assert res.confidence >= 0.75
assert res.needs_quarantine is False
def test_classify_movie_with_apple_tv_tag(classifier):
scanned = ScannedFile(path=Path("/downloads/Wolfs.2024.1080p.Apple.TV.WEB-DL.DDP5.1.Atmos.H.264.mkv"), size=4500000000, mtime=1000.0)
tokens = TokenizedFilename(
raw_name="Wolfs.2024.1080p.Apple.TV.WEB-DL.DDP5.1.Atmos.H.264.mkv",
title="Wolfs",
year=2024,
resolution="1080p",
video_codec="H.264",
source="WEB-DL",
)
meta = MediaMetadata(
path=scanned.path,
mime_type="video/x-matroska",
container="mkv",
duration_seconds=6400,
has_video=True,
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "movie"
assert res.confidence >= 0.75
def test_video_file_with_audio_not_classified_as_music(classifier):
# Video container .mkv with audio track should never be classified as music
scanned = ScannedFile(
path=Path("/downloads/Star.Wars.The.Clone.Wars.S01E01.1080p.BluRay.REMUX.VC-1.DD5.1-NOGRP.mkv"),
size=4500000000,
mtime=1000.0,
)
tokens = TokenizedFilename(
raw_name="Star.Wars.The.Clone.Wars.S01E01.1080p.BluRay.REMUX.VC-1.DD5.1-NOGRP.mkv",
title="Star Wars The Clone Wars",
season=1,
episode=1,
is_episodic=True,
)
meta = MediaMetadata(
path=scanned.path,
mime_type="video/x-matroska",
container="mkv",
has_audio=True,
has_video=False, # e.g. exotic codec in container
)
res = classifier.classify(scanned, tokens, meta)
assert res.category == "tv"
assert res.confidence >= 0.80
assert res.needs_quarantine is False
+75
View File
@@ -0,0 +1,75 @@
import os
import pytest
from pathlib import Path
from media_sorter.config import Settings, ActionType, ConflictPolicy
from media_sorter.db import get_engine, init_db, get_db_session
from media_sorter.models import BatchRecord, Operation, FileRecord, QuarantineRecord
def test_settings_defaults():
settings = Settings()
assert settings.general.dry_run is True
assert settings.general.confidence_threshold == 0.75
assert settings.general.action == ActionType.MOVE
assert settings.conflicts.policy == ConflictPolicy.RENAME_UNIQUE
assert settings.conflicts.allow_overwrite is False
assert "*.txt" in settings.filters.exclude_patterns
from media_sorter.scanner import Scanner
scanner = Scanner()
assert "*.txt" in scanner.exclude_patterns
assert not scanner._matches_filter("info.txt")
assert not scanner._matches_filter("README.TXT")
assert scanner._matches_filter("movie.mkv")
def test_load_yaml_config(tmp_path: Path):
config_file = tmp_path / "test_config.yaml"
config_file.write_text("""
general:
dry_run: false
confidence_threshold: 0.85
storage:
source_dirs:
- "/test/source"
destination_base: "/test/organized"
""", encoding="utf-8")
settings = Settings.load_from_file(config_file)
assert settings.general.dry_run is False
assert settings.general.confidence_threshold == 0.85
assert "/test/source" in settings.storage.source_dirs
def test_database_initialization(tmp_path: Path):
db_file = tmp_path / "test.db"
engine = init_db(db_path=db_file)
assert db_file.exists()
with get_db_session(engine) as session:
batch = BatchRecord(id="test-uuid-1", dry_run=True, status="COMPLETED")
session.add(batch)
with get_db_session(engine) as session:
queried = session.query(BatchRecord).filter_by(id="test-uuid-1").first()
assert queried is not None
assert queried.dry_run is True
assert queried.status == "COMPLETED"
def test_session_factory_caching_and_cleanup(tmp_path: Path):
from media_sorter.db import get_session_factory, _ENGINE_SESSION_FACTORIES
db_file = tmp_path / "test_cache.db"
engine = init_db(db_path=db_file)
factory1 = get_session_factory(engine)
factory2 = get_session_factory(engine)
assert factory1 is factory2
assert engine in _ENGINE_SESSION_FACTORIES
with get_db_session(engine) as session:
batch = BatchRecord(id="cached-1", dry_run=True, status="COMPLETED")
session.add(batch)
# Scoped session registry was cleared via remove()
assert factory1.registry.has() is False
+268
View File
@@ -0,0 +1,268 @@
from pathlib import Path
import pytest
from media_sorter.config import ActionType, ConflictPolicy, Settings
from media_sorter.db import get_db_session, init_db
from media_sorter.executor import MediaExecutor, PlannedOperation
from media_sorter.models import BatchRecord, Operation, OperationStatus
@pytest.fixture
def temp_env(tmp_path: Path):
db_path = tmp_path / "test.db"
engine = init_db(db_path=db_path)
src_dir = tmp_path / "incoming"
src_dir.mkdir()
dst_dir = tmp_path / "organized"
dst_dir.mkdir()
settings = Settings()
settings.database.path = str(db_path)
settings.storage.destination_base = str(dst_dir)
settings.storage.source_dirs = [str(src_dir)]
return settings, engine, src_dir, dst_dir
def test_dry_run_leaves_filesystem_untouched(temp_env):
settings, engine, src_dir, dst_dir = temp_env
test_file = src_dir / "sample_movie.mkv"
test_file.write_text("dummy content")
target_dst = dst_dir / "Movies/sample_movie.mkv"
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
plan = [
PlannedOperation(
src=test_file,
dst=target_dst,
action=ActionType.MOVE,
category="movie",
confidence=0.95,
)
]
report = executor.execute_batch(plan, dry_run=True)
assert report.dry_run is True
assert report.moved_files == 1
# File must still exist at src and not at dst
assert test_file.exists()
assert not target_dst.exists()
def test_live_atomic_move_and_rollback(temp_env):
settings, engine, src_dir, dst_dir = temp_env
settings.general.dry_run = False
test_file = src_dir / "song.flac"
test_file.write_text("flac audio bytes")
target_dst = dst_dir / "Music/Artist/Album/01 - song.flac"
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
plan = [
PlannedOperation(
src=test_file,
dst=target_dst,
action=ActionType.MOVE,
category="music",
confidence=0.95,
)
]
report = executor.execute_batch(plan, dry_run=False)
assert report.moved_files == 1
assert not test_file.exists()
assert target_dst.exists()
assert target_dst.read_text() == "flac audio bytes"
# Now execute rollback
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
reverted = executor.rollback_batch(report.batch_id)
assert reverted == 1
assert test_file.exists()
assert not target_dst.exists()
assert test_file.read_text() == "flac audio bytes"
def test_conflict_rename_unique(temp_env):
settings, engine, src_dir, dst_dir = temp_env
settings.general.dry_run = False
settings.conflicts.policy = ConflictPolicy.RENAME_UNIQUE
# Pre-create existing file at destination
target_dst = dst_dir / "Movies/Avatar (2009)/Avatar (2009).mkv"
target_dst.parent.mkdir(parents=True, exist_ok=True)
target_dst.write_text("original 1080p copy")
# New incoming file
incoming = src_dir / "Avatar.2009.2160p.mkv"
incoming.write_text("new 4k copy")
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
plan = [
PlannedOperation(
src=incoming,
dst=target_dst,
action=ActionType.MOVE,
category="movie",
confidence=0.95,
)
]
report = executor.execute_batch(plan, dry_run=False)
assert report.moved_files == 1
assert target_dst.exists()
assert target_dst.read_text() == "original 1080p copy"
# New file should have been renamed uniquely: Avatar (2009) (1).mkv
unique_dst = dst_dir / "Movies/Avatar (2009)/Avatar (2009) (1).mkv"
assert unique_dst.exists()
assert unique_dst.read_text() == "new 4k copy"
def test_conflict_replace_with_backup_and_rollback(temp_env):
settings, engine, src_dir, dst_dir = temp_env
settings.general.dry_run = False
settings.conflicts.policy = ConflictPolicy.REPLACE_IF_HIGHER_QUALITY
backup_dir = src_dir.parent / ".backup"
settings.conflicts.backup_dir = str(backup_dir)
target_dst = dst_dir / "Movies/Test.mkv"
target_dst.parent.mkdir(parents=True, exist_ok=True)
target_dst.write_text("old version")
incoming = src_dir / "Test.mkv"
incoming.write_text("upgraded high quality version")
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
plan = [
PlannedOperation(
src=incoming,
dst=target_dst,
action=ActionType.MOVE,
category="movie",
confidence=0.95,
)
]
report = executor.execute_batch(plan, dry_run=False)
assert target_dst.read_text() == "upgraded high quality version"
# Rollback should restore the old version from backup!
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
reverted = executor.rollback_batch(report.batch_id)
assert reverted == 1
assert target_dst.read_text() == "old version"
assert incoming.read_text() == "upgraded high quality version"
def test_process_locking(tmp_path: Path):
from media_sorter.executor import acquire_process_lock, ProcessLockError
lock_file = tmp_path / "test.lock"
with acquire_process_lock(lock_file):
# Trying to acquire same lock file concurrently must fail with ProcessLockError
with pytest.raises(ProcessLockError):
with acquire_process_lock(lock_file):
pass
# After exiting the lock block, it should be cleanly re-acquirable
with acquire_process_lock(lock_file):
pass
def test_cleanup_deletes_txt_files_and_removes_dirs(temp_env):
settings, engine, src_dir, dst_dir = temp_env
settings.general.dry_run = False
settings.general.cleanup_empty_dirs = True
# Setup source subdirectory with a movie, a .txt file, and a companion .txt file
sub_dir = src_dir / "Movie.Release.2023"
sub_dir.mkdir(parents=True, exist_ok=True)
movie_file = sub_dir / "movie.mkv"
movie_file.write_text("dummy video")
companion_txt = sub_dir / "movie.txt"
companion_txt.write_text("companion text")
readme_txt = sub_dir / "README.txt"
readme_txt.write_text("torrent info")
root_txt = src_dir / "root_note.txt"
root_txt.write_text("root text file")
target_dst = dst_dir / "Movies/Movie (2023)/movie.mkv"
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
plan = [
PlannedOperation(
src=movie_file,
dst=target_dst,
action=ActionType.MOVE,
category="movie",
confidence=0.95,
)
]
report = executor.execute_batch(plan, dry_run=False)
assert report.moved_files == 1
assert target_dst.exists()
# Verify movie is gone from src
assert not movie_file.exists()
# Verify .txt files were deleted during cleanup
assert not companion_txt.exists()
assert not readme_txt.exists()
assert not root_txt.exists()
# Verify the empty subdirectory was cleaned up (rmdir'd)
assert not sub_dir.exists()
# Verify source root itself was NOT removed
assert src_dir.exists()
def test_rollback_all_batches(temp_env):
settings, engine, src_dir, dst_dir = temp_env
settings.general.dry_run = False
file1 = src_dir / "sample1.mkv"
file1.write_text("file 1")
file2 = src_dir / "sample2.mkv"
file2.write_text("file 2")
dst1 = dst_dir / "Movies/Movie 1/sample1.mkv"
dst2 = dst_dir / "Movies/Movie 2/sample2.mkv"
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
executor.execute_batch(
[PlannedOperation(src=file1, dst=dst1, action=ActionType.MOVE, category="movie", confidence=0.9)],
dry_run=False
)
executor.execute_batch(
[PlannedOperation(src=file2, dst=dst2, action=ActionType.MOVE, category="movie", confidence=0.9)],
dry_run=False
)
assert not file1.exists()
assert not file2.exists()
assert dst1.exists()
assert dst2.exists()
with get_db_session(engine) as session:
executor = MediaExecutor(settings, session)
reverted = executor.rollback_all()
assert reverted == 2
assert file1.exists()
assert file2.exists()
assert not dst1.exists()
assert not dst2.exists()
+253
View File
@@ -0,0 +1,253 @@
from pathlib import Path
import pytest
from fastapi.testclient import TestClient
from media_sorter.config import Settings
from media_sorter.db import get_db_session, init_db
from media_sorter.library import (
clean_show_title,
get_known_shows,
list_library_items,
match_known_show,
record_detected_item,
sync_library_from_disk,
)
from media_sorter.server import cluster_unsure_files, create_app, inspect_downloads_folder
@pytest.fixture
def library_env(tmp_path: Path):
downloads = tmp_path / "downloads"
movies = tmp_path / "movies"
shows = tmp_path / "shows"
db_file = tmp_path / "test.db"
downloads.mkdir()
movies.mkdir()
shows.mkdir()
settings = Settings()
settings.storage.source_dirs = [str(downloads)]
settings.storage.destination_dirs.movies = str(movies)
settings.storage.destination_dirs.tv = str(shows)
settings.database.path = str(db_file)
settings.general.dry_run = False
settings.general.min_file_age_seconds = 0
engine = init_db(db_path=db_file)
test_env_file = tmp_path / ".env"
app = create_app(settings, engine, env_path=test_env_file)
client = TestClient(app)
return client, settings, engine, downloads, movies, shows
def test_library_sync_and_show_memory(library_env):
client, settings, engine, downloads, movies, shows = library_env
# 1. Populate disk with shows and movies
dexter_dir = shows / "Dexter" / "Season 01"
dexter_dir.mkdir(parents=True)
(dexter_dir / "Dexter - S01E01.mkv").write_bytes(b"\x00" * 100)
(dexter_dir / "Dexter - S01E02.mkv").write_bytes(b"\x00" * 100)
breaking_bad_dir = shows / "Breaking Bad" / "Season 01"
breaking_bad_dir.mkdir(parents=True)
(breaking_bad_dir / "Breaking Bad - S01E01.mkv").write_bytes(b"\x00" * 100)
movie_dir = movies / "Inception (2010)"
movie_dir.mkdir(parents=True)
(movie_dir / "Inception (2010).mkv").write_bytes(b"\x00" * 100)
with get_db_session(engine) as session:
sync_res = sync_library_from_disk(session, settings)
assert sync_res["shows_synced"] == 2
assert sync_res["movies_synced"] == 1
# 2. Check list_library_items
all_items = list_library_items(session)
assert all_items["total_shows"] == 2
assert all_items["total_movies"] == 1
shows_only = list_library_items(session, category="tv")
assert len(shows_only["shows"]) == 2
assert len(shows_only["movies"]) == 0
movies_only = list_library_items(session, category="movie")
assert len(movies_only["movies"]) == 1
search_res = list_library_items(session, search="dexter")
assert len(search_res["shows"]) == 1
assert search_res["shows"][0]["title"] == "Dexter"
# 3. Test known shows matching
known_shows = get_known_shows(session)
assert len(known_shows) == 2
matched = match_known_show("Dexter's Kill Room Extra.mkv", known_shows)
assert matched is not None
assert matched["title"] == "Dexter"
matched_bb = match_known_show("Breaking.Bad.Behind.The.Scenes.mp4", known_shows)
assert matched_bb is not None
assert matched_bb["title"] == "Breaking Bad"
assert match_known_show("Unrelated Movie.mkv", known_shows) is None
def test_cluster_unsure_files(library_env):
client, settings, engine, downloads, movies, shows = library_env
unsure_files = [
# Subfolder group (Folder: Dexter Extras)
{"name": "Interview.mkv", "relative_path": "Dexter Extras/Interview.mkv", "detected_type": "other"},
{"name": "Behind Scenes.mkv", "relative_path": "Dexter Extras/Behind Scenes.mkv", "detected_type": "other"},
# Common prefix group (Blood, Guts and Body Parts)
{"name": "Blood, Guts and Body Parts The Blood.mkv", "relative_path": "Blood, Guts and Body Parts The Blood.mkv", "detected_type": "other"},
{"name": "Blood, Guts and Body Parts The Props.mkv", "relative_path": "Blood, Guts and Body Parts The Props.mkv", "detected_type": "other"},
# Same title group (Inception)
{"name": "Inception.1080p.mkv", "relative_path": "Inception.1080p.mkv", "detected_type": "movie"},
{"name": "Inception.720p.mp4", "relative_path": "Inception.720p.mp4", "detected_type": "movie"},
# Standalone single file
{"name": "Solo Movie (2021).mkv", "relative_path": "Solo Movie (2021).mkv", "detected_type": "movie"},
]
groups, singles = cluster_unsure_files(unsure_files, settings)
group_names = [g["group_name"] for g in groups]
assert "Dexter Extras" in group_names
assert any("Blood" in gn for gn in group_names)
single_names = [s["name"] for s in singles]
assert "Solo Movie (2021).mkv" in single_names
def test_show_memory_routing_in_inspection(library_env):
client, settings, engine, downloads, movies, shows = library_env
# Record "Dexter" into library
with get_db_session(engine) as session:
record_detected_item(session, settings, "Dexter", "tv")
# Add a non-standard extra in downloads matching Dexter
dexter_extra = downloads / "Dexter's Kill Room Extra.mkv"
dexter_extra.write_bytes(b"\x00" * 100)
# Inspect downloads with engine
inspection = inspect_downloads_folder(downloads, settings, engine=engine)
assert len(inspection["shows"]) == 1
assert inspection["shows"][0]["show_name"] == "Dexter"
assert len(inspection["shows"][0]["files"]) == 1
def test_library_and_sort_group_api(library_env):
client, settings, engine, downloads, movies, shows = library_env
# 1. API: Rescan Library
dexter_dir = shows / "Dexter" / "Season 01"
dexter_dir.mkdir(parents=True)
(dexter_dir / "Dexter - S01E01.mkv").write_bytes(b"\x00" * 100)
rescan_resp = client.post("/api/library/rescan")
assert rescan_resp.status_code == 200
assert rescan_resp.json()["shows_synced"] >= 1
# 2. API: GET /api/library
lib_resp = client.get("/api/library")
assert lib_resp.status_code == 200
lib_data = lib_resp.json()
assert lib_data["total_shows"] >= 1
# 3. API: POST /api/files/sort-group
group_sub = downloads / "My Special Show"
group_sub.mkdir()
f1 = group_sub / "Episode 1.mkv"
f2 = group_sub / "Episode 2.mkv"
f1.write_bytes(b"\x00" * 100)
f2.write_bytes(b"\x00" * 100)
sort_group_resp = client.post("/api/files/sort-group", json={
"group_name": "My Special Show",
"group_type": "folder",
"category": "tv",
"title": "My Special Show",
"relative_paths": ["My Special Show/Episode 1.mkv", "My Special Show/Episode 2.mkv"],
})
assert sort_group_resp.status_code == 200
sort_data = sort_group_resp.json()
assert sort_data["status"] == "ok"
assert sort_data["moved_files"] == 2
# Verify files moved to shows directory
dest_show = shows / "My Special Show"
assert dest_show.exists()
# Verify newly sorted show was automatically recorded in the library
lib_after = client.get("/api/library?search=Special").json()
assert len(lib_after["shows"]) == 1
assert lib_after["shows"][0]["title"] == "My Special Show"
def test_one_piece_batch_sorting_and_downloads_inspection(library_env):
client, settings, engine, downloads, movies, shows = library_env
# 1. Create simulated One Piece files in downloads (unordered to test natural sorting)
s8 = downloads / "One Piece (0001-1071+Movies+Specials)" / "Season 08 - Water Seven (229-263)"
s8.mkdir(parents=True)
f233 = s8 / "[Anime Time] One Piece - 0233 - Pirate Abduction Incident!.mkv"
f232 = s8 / "[Anime Time] One Piece - 0232 - Galley-La Company!.mkv"
f1171 = downloads / "One.Piece.S01E1171.1080p.CR.WEB-DL.mkv"
f1073 = downloads / "[HatSubs] One Piece 1073 (BD 1080p 10-bit Opus) v2 [194B3FBA].mkv"
f233.write_bytes(b"\x00" * 100)
f232.write_bytes(b"\x00" * 100)
f1171.write_bytes(b"\x00" * 100)
f1073.write_bytes(b"\x00" * 100)
# 2. Inspect downloads folder via /api/files
files_res = client.get("/api/files").json()
dl_data = files_res["downloads"]
assert len(dl_data["shows"]) == 1
op_show = dl_data["shows"][0]
assert op_show["show_name"] == "One Piece"
assert op_show["count"] == 4
# Verify all episodes extracted and ordered naturally (232 before 233)
ep_map = {f["name"]: (f["season"], f["episode"]) for f in op_show["files"]}
assert ep_map["[Anime Time] One Piece - 0232 - Galley-La Company!.mkv"] == (8, 232)
assert ep_map["[Anime Time] One Piece - 0233 - Pirate Abduction Incident!.mkv"] == (8, 233)
assert ep_map["One.Piece.S01E1171.1080p.CR.WEB-DL.mkv"] == (1, 1171)
assert ep_map["[HatSubs] One Piece 1073 (BD 1080p 10-bit Opus) v2 [194B3FBA].mkv"] == (1, 1073)
# Verify natural order: 232 appears before 233
idx_232 = next(i for i, f in enumerate(op_show["files"]) if f["episode"] == 232)
idx_233 = next(i for i, f in enumerate(op_show["files"]) if f["episode"] == 233)
assert idx_232 < idx_233
# 3. Call manual sort batch API
payload = {
"show_name": "One Piece",
"files": [
{"relative_path": str(f232.relative_to(downloads)), "season": 8, "episode": 232},
{"relative_path": str(f233.relative_to(downloads)), "season": 8, "episode": 233},
{"relative_path": str(f1171.relative_to(downloads)), "season": 1, "episode": 1171},
{"relative_path": str(f1073.relative_to(downloads)), "season": 1, "episode": 1073},
]
}
batch_res = client.post("/api/files/manual-sort-batch", json=payload)
assert batch_res.status_code == 200
assert batch_res.json()["status"] == "ok"
assert batch_res.json()["moved_files"] == 4
# 4. Check destinations exist and are correctly named
dest_232 = shows / "One Piece" / "Season 08" / "One Piece - S08E232.mkv"
dest_233 = shows / "One Piece" / "Season 08" / "One Piece - S08E233.mkv"
dest_1073 = shows / "One Piece" / "Season 01" / "One Piece - S01E1073.mkv"
dest_1171 = shows / "One Piece" / "Season 01" / "One Piece - S01E1171.mkv"
assert dest_232.exists(), "Episode 232 must exist in Season 08"
assert dest_233.exists(), "Episode 233 must exist in Season 08"
assert dest_1073.exists(), "Episode 1073 must exist in Season 01"
assert dest_1171.exists(), "Episode 1171 must exist in Season 01"
# 5. Check library records One Piece show
lib = client.get("/api/library?search=One+Piece").json()
assert len(lib["shows"]) == 1
assert lib["shows"][0]["title"] == "One Piece"
+94
View File
@@ -0,0 +1,94 @@
from pathlib import Path
import pytest
from media_sorter.analyzer import MediaMetadata
from media_sorter.classifier import ClassificationResult
from media_sorter.config import Settings
from media_sorter.namer import MediaNamer, sanitize_filename_component
from media_sorter.tokenizer import TokenizedFilename
@pytest.fixture
def settings():
return Settings()
@pytest.fixture
def namer(settings):
return MediaNamer(settings)
def test_sanitize_filename_forbidden_chars():
messy = 'Movie: "The Final Chapter" <Director\'s Cut> | Part 1?.mkv'
cleaned = sanitize_filename_component(messy)
assert ":" not in cleaned
assert '"' not in cleaned
assert "<" not in cleaned
assert ">" not in cleaned
assert "|" not in cleaned
assert "?" not in cleaned
assert cleaned.endswith(".mkv")
def test_sanitize_windows_reserved_names():
res = sanitize_filename_component("CON.mp4")
assert res == "_CON.mp4"
res2 = sanitize_filename_component("nul.txt")
assert res2 == "_nul.txt"
def test_generate_movie_destination(namer):
tokens = TokenizedFilename(
raw_name="The.Matrix.1999.1080p.BluRay.x264.mkv",
title="The Matrix",
year=1999,
resolution="1080p",
video_codec="x264",
)
meta = MediaMetadata(path=Path("The.Matrix.1999.1080p.BluRay.x264.mkv"), mime_type="video/x-matroska", container="mkv")
cls_res = ClassificationResult(category="movie", confidence=0.95, tokens=tokens, metadata=meta)
dest = namer.generate_destination_path(cls_res)
assert "The Matrix (1999)" in str(dest)
assert dest.suffix == ".mkv"
def test_generate_tv_destination(namer):
tokens = TokenizedFilename(
raw_name="Breaking Bad S01E01 Pilot.mkv",
title="Breaking Bad",
season=1,
episode=1,
episode_title="Pilot",
)
meta = MediaMetadata(path=Path("Breaking Bad S01E01 Pilot.mkv"), mime_type="video/x-matroska", container="mkv")
cls_res = ClassificationResult(category="tv", confidence=0.95, tokens=tokens, metadata=meta)
dest = namer.generate_destination_path(cls_res)
assert "Season 01" in str(dest)
assert "Breaking Bad" in str(dest)
assert "S01E01" in str(dest)
def test_generate_quarantine_destination(namer):
meta = MediaMetadata(path=Path("weird_unknown_file.xyz"), mime_type="application/octet-stream", container="xyz")
cls_res = ClassificationResult(
category="unknown",
confidence=0.1,
metadata=meta,
needs_quarantine=True,
quarantine_reason="unrecognized_format",
)
dest = namer.generate_destination_path(cls_res)
assert "Quarantine" in str(dest)
assert "unrecognized-format" in str(dest) or "unrecognized_format" in str(dest)
def test_sidecar_subtitle_matching(namer):
sub_meta = MediaMetadata(path=Path("movie.en.srt"), mime_type="text/plain", container="srt")
cls_res = ClassificationResult(category="subtitle", confidence=0.95, metadata=sub_meta)
primary_dst = Path("/organized/Movies/Inception (2010)/Inception (2010) [1080p].mkv")
sub_dst = namer.generate_destination_path(cls_res, primary_dst_path=primary_dst)
assert sub_dst.parent == primary_dst.parent
assert sub_dst.name == "Inception (2010) [1080p].en.srt"
+112
View File
@@ -0,0 +1,112 @@
from pathlib import Path
import pytest
from media_sorter.db import get_db_session, init_db
from media_sorter.models import QuarantineRecord, QuarantineStatus
from media_sorter.quarantine import QuarantineManager
@pytest.fixture
def session(tmp_path: Path):
db_file = tmp_path / "test.db"
engine = init_db(db_path=db_file)
with get_db_session(engine) as s:
yield s
def test_quarantine_crud(session):
qm = QuarantineManager(session)
# Insert pending item
item = QuarantineRecord(
src="/incoming/ambiguous.avi",
suggested_category="movie",
confidence=0.55,
reason="Missing release year and technical tags",
status=QuarantineStatus.PENDING.value,
)
session.add(item)
session.commit()
pending = qm.list_pending()
assert len(pending) == 1
assert pending[0].src == "/incoming/ambiguous.avi"
# Resolve item
success = qm.resolve_item(pending[0].id, "tv", "/organized/TV/Show/ep.avi")
assert success is True
resolved = qm.get_by_id(pending[0].id)
assert resolved.status == QuarantineStatus.RESOLVED.value
assert resolved.suggested_category == "tv"
assert resolved.resolved_path == "/organized/TV/Show/ep.avi"
# Pending list should now be empty
assert len(qm.list_pending()) == 0
stats = qm.get_statistics()
assert stats["total"] == 1
assert stats["pending"] == 0
assert stats["resolved"] == 1
def test_quarantine_undo(session, tmp_path: Path):
qm = QuarantineManager(session)
# 1. Test unflagging a pending item
src_file = tmp_path / "pending_sample.mkv"
src_file.write_text("dummy")
item1 = QuarantineRecord(
src=str(src_file),
suggested_category="movie",
confidence=0.50,
reason="Low confidence",
status=QuarantineStatus.PENDING.value,
)
session.add(item1)
session.commit()
assert len(qm.list_pending()) == 1
# Undo pending item should delete/unflag it
success = qm.undo_item(item1.id)
assert success is True
assert len(qm.list_pending()) == 0
assert qm.get_by_id(item1.id) is None
# 2. Test undoing a resolved item (moves file back to src)
if src_file.exists():
src_file.unlink()
dst_file = tmp_path / "organized" / "Movie (2020)" / "Movie (2020).mkv"
dst_file.parent.mkdir(parents=True, exist_ok=True)
dst_file.write_text("movie data")
item2 = QuarantineRecord(
src=str(src_file),
suggested_category="movie",
confidence=0.60,
reason="Ambiguous",
status=QuarantineStatus.RESOLVED.value,
resolved_path=str(dst_file),
)
session.add(item2)
session.commit()
assert len(qm.list_resolved()) == 1
assert dst_file.exists()
assert not src_file.exists()
success2 = qm.undo_item(item2.id)
assert success2 is True
# File should be moved back to src
assert src_file.exists()
assert src_file.read_text() == "movie data"
assert not dst_file.exists()
# Record should now be back to PENDING with resolved fields reset
rec2 = qm.get_by_id(item2.id)
assert rec2.status == QuarantineStatus.PENDING.value
assert rec2.resolved_path is None
assert len(qm.list_pending()) == 1
assert len(qm.list_resolved()) == 0
+201
View File
@@ -0,0 +1,201 @@
import os
from pathlib import Path
import socket
import pytest
from fastapi.testclient import TestClient
from media_sorter.config import Settings
from media_sorter.db import get_db_session, init_db
from media_sorter.library import record_detected_item
from media_sorter.models import LibraryItem
from media_sorter.server import create_app
@pytest.fixture
def isolated_web_env(tmp_path: Path):
"""Isolated environment with temporary directories and SQLite database."""
downloads = tmp_path / "downloads"
movies = tmp_path / "movies"
shows = tmp_path / "shows"
db_file = tmp_path / "test.db"
downloads.mkdir()
movies.mkdir()
shows.mkdir()
settings = Settings()
settings.storage.source_dirs = [str(downloads)]
settings.storage.destination_dirs.movies = str(movies)
settings.storage.destination_dirs.tv = str(shows)
settings.database.path = str(db_file)
settings.general.dry_run = False
settings.general.min_file_age_seconds = 0
engine = init_db(db_path=db_file)
test_env_file = tmp_path / ".env"
app = create_app(settings, engine, env_path=test_env_file)
client = TestClient(app)
return client, settings, engine, downloads, movies, shows, test_env_file
# -----------------------------------------------------------------------------
# Test 1: Production Filesystem Safety Trap
# -----------------------------------------------------------------------------
def test_protect_production_filesystem_triggers_on_write(tmp_path: Path):
"""Verify protect_production_filesystem trap triggers RuntimeError on attempted
write or delete in /md0/jdownloads/illegal.txt, /md0/movies1, /md0/tv1.
"""
illegal_download = Path("/md0/jdownloads/illegal.txt")
with pytest.raises(RuntimeError, match="FILESYSTEM SAFETY TRAP"):
illegal_download.write_text("dangerous write")
with pytest.raises(RuntimeError, match="FILESYSTEM SAFETY TRAP"):
illegal_download.write_bytes(b"dangerous bytes")
with pytest.raises(RuntimeError, match="FILESYSTEM SAFETY TRAP"):
illegal_download.unlink()
with pytest.raises(RuntimeError, match="FILESYSTEM SAFETY TRAP"):
os.remove("/md0/jdownloads/illegal.txt")
with pytest.raises(RuntimeError, match="FILESYSTEM SAFETY TRAP"):
open("/md0/jdownloads/illegal.txt", "w")
# Verify protection for /md0/movies1 and /md0/tv1
with pytest.raises(RuntimeError, match="FILESYSTEM SAFETY TRAP"):
Path("/md0/movies1/illegal_movie.mkv").touch()
with pytest.raises(RuntimeError, match="FILESYSTEM SAFETY TRAP"):
os.makedirs("/md0/tv1/illegal_show/Season 01")
# Verify tmp_path operations succeed normally
safe_file = tmp_path / "safe.txt"
safe_file.write_text("allowed content")
assert safe_file.read_text() == "allowed content"
safe_file.unlink()
assert not safe_file.exists()
# -----------------------------------------------------------------------------
# Test 2: Test Environment Isolation
# -----------------------------------------------------------------------------
def test_isolate_test_environment_step_1_mutates_environment():
"""Step 1: mutate CONFIDENCE_THRESHOLD in os.environ and confirm Settings uses it."""
os.environ["CONFIDENCE_THRESHOLD"] = "0.99"
settings = Settings()
assert settings.general.confidence_threshold == 0.99
def test_isolate_test_environment_step_2_reverts_to_clean_defaults():
"""Step 2: verify isolate_test_environment restored os.environ and clean defaults."""
assert os.environ.get("CONFIDENCE_THRESHOLD") != "0.99"
assert "CONFIDENCE_THRESHOLD" not in os.environ
settings = Settings()
assert settings.general.confidence_threshold == 0.75
assert settings.general.worker_count == 4
# -----------------------------------------------------------------------------
# Test 3: /api/explorer/set-destination Latent NameError Fix
# -----------------------------------------------------------------------------
def test_explorer_set_destination_does_not_raise_name_error(isolated_web_env):
"""Verify /api/explorer/set-destination returns 200 without NameError."""
client, settings, engine, downloads, movies, shows, _ = isolated_web_env
resp = client.post(
"/api/explorer/set-destination",
json={"show_idx": 0, "destination": "/custom/destination/folder"},
)
assert resp.status_code == 200
data = resp.json()
assert data.get("status") == "ok"
assert data.get("destination") == "/custom/destination/folder"
assert data.get("success") is True
# Check validation for missing parameters
bad_resp = client.post("/api/explorer/set-destination", json={"show_idx": 0})
assert bad_resp.status_code == 400
# -----------------------------------------------------------------------------
# Test 4: Read-Only File Inspection Does Not Inflate item_count
# -----------------------------------------------------------------------------
def test_get_files_does_not_inflate_library_item_count(isolated_web_env):
"""Verify GET /api/files does not inflate LibraryItem.item_count on repeated calls."""
client, settings, engine, downloads, movies, shows, _ = isolated_web_env
# 1. Place a show episode in downloads
test_file = downloads / "Breaking.Bad.S01E01.1080p.mkv"
test_file.write_text("media bytes")
# Seed an existing library item with initial count 5
with get_db_session(engine) as session:
record_detected_item(
session,
settings,
"Breaking Bad",
"tv",
destination_folder=str(shows / "Breaking Bad"),
delta_count=5,
)
# Query before calling /api/files
with get_db_session(engine) as session:
item_before = (
session.query(LibraryItem).filter_by(title="Breaking Bad", category="tv").first()
)
assert item_before is not None
assert item_before.item_count == 5
# 2. Call GET /api/files repeatedly (5 times)
for _ in range(5):
resp = client.get("/api/files")
assert resp.status_code == 200
# 3. Verify item_count remained 5 and was NOT inflated
with get_db_session(engine) as session:
item_after = (
session.query(LibraryItem).filter_by(title="Breaking Bad", category="tv").first()
)
assert item_after is not None
assert item_after.item_count == 5
# -----------------------------------------------------------------------------
# Test 5: /api/files/scan Alias Routes (GET & POST)
# -----------------------------------------------------------------------------
def test_api_files_scan_alias_get_and_post(isolated_web_env):
"""Verify GET /api/files/scan and POST /api/files/scan return HTTP 200 with
identical data to GET /api/files.
"""
client, settings, engine, downloads, movies, shows, _ = isolated_web_env
(downloads / "Severance.S01E01.mkv").write_text("dummy video")
(movies / "Inception (2010)").mkdir(parents=True, exist_ok=True)
(movies / "Inception (2010)" / "Inception (2010).mkv").write_text("dummy movie")
resp_base = client.get("/api/files")
assert resp_base.status_code == 200
base_data = resp_base.json()
resp_scan_get = client.get("/api/files/scan")
assert resp_scan_get.status_code == 200
assert resp_scan_get.json() == base_data
resp_scan_post = client.post("/api/files/scan")
assert resp_scan_post.status_code == 200
assert resp_scan_post.json() == base_data
# Check top-level contract keys
assert "downloads" in base_data
assert "movies" in base_data
assert "shows" in base_data
# -----------------------------------------------------------------------------
# Bonus Test: External Network Trap
# -----------------------------------------------------------------------------
def test_block_external_network_trap():
"""Verify block_external_network prevents outbound external connections."""
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
with pytest.raises(RuntimeError, match="NETWORK ACCESS TRAP"):
s.connect(("8.8.8.8", 53))
+733
View File
@@ -0,0 +1,733 @@
from pathlib import Path
import pytest
from fastapi.testclient import TestClient
from media_sorter.config import Settings
from media_sorter.db import init_db
from media_sorter.server import create_app
@pytest.fixture
def web_env(tmp_path: Path):
downloads = tmp_path / "downloads"
movies = tmp_path / "movies"
shows = tmp_path / "shows"
db_file = tmp_path / "test.db"
downloads.mkdir()
movies.mkdir()
shows.mkdir()
settings = Settings()
settings.storage.source_dirs = [str(downloads)]
settings.storage.destination_dirs.movies = str(movies)
settings.storage.destination_dirs.tv = str(shows)
settings.database.path = str(db_file)
settings.general.dry_run = False
settings.general.min_file_age_seconds = 0
engine = init_db(db_path=db_file)
test_env_file = tmp_path / ".env"
app = create_app(settings, engine, env_path=test_env_file)
client = TestClient(app)
return client, settings, downloads, movies, shows, test_env_file
def test_web_status_and_dashboard(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# 1. HTML Dashboard renders cleanly
resp = client.get("/")
assert resp.status_code == 200
assert "Media Sorter" in resp.text
assert "Folder Explorer" in resp.text
# 2. Status API returns paths
status_resp = client.get("/api/status")
assert status_resp.status_code == 200
data = status_resp.json()
assert data["status"] == "online"
assert data["downloads_dir"] == str(downloads)
assert data["movies_dir"] == str(movies)
assert data["shows_dir"] == str(shows)
def test_web_sample_generation_and_sorting(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# 1. Generate sample downloads
sample_resp = client.post("/api/files/test-sample")
assert sample_resp.status_code == 200
sample_data = sample_resp.json()
assert len(sample_data["files"]) >= 5
# Verify files in downloads folder via API
files_resp = client.get("/api/files")
assert files_resp.status_code == 200
files_data = files_resp.json()
assert len(files_data["downloads"]["files"]) >= 5
# 2. Trigger Sorter Live execution via Web API
run_resp = client.post("/api/run", json={"dry_run": False})
assert run_resp.status_code == 200
run_data = run_resp.json()
assert run_data["moved_files"] >= 3
# Verify files moved to movies and shows
files_after = client.get("/api/files").json()
assert len(files_after["movies"]["files"]) >= 2
assert len(files_after["shows"]["files"]) >= 2
# 3. Trigger Rollback via Web API
rb_resp = client.post("/api/rollback", json={"batch_id": run_data["batch_id"]})
assert rb_resp.status_code == 200
rb_data = rb_resp.json()
assert rb_data["reverted_files"] >= 3
# Verify files restored to downloads folder
files_reverted = client.get("/api/files").json()
assert len(files_reverted["downloads"]["files"]) >= 4
# 4. Test deleting a single download file via Web API
file_to_del = sample_data["files"][0]
del_resp = client.delete(f"/api/files/download?name={file_to_del}")
assert del_resp.status_code == 200
assert del_resp.json()["status"] == "deleted"
files_post_del = client.get("/api/files").json()
assert len(files_post_del["downloads"]["files"]) == len(files_reverted["downloads"]["files"]) - 1
def test_web_settings_update(web_env, tmp_path: Path):
client, settings, downloads, movies, shows, test_env_file = web_env
new_downloads = tmp_path / "new_downloads"
new_movies = tmp_path / "new_movies"
update_payload = {
"downloads_dir": str(new_downloads),
"movies_dir": str(new_movies),
"dry_run": True,
"confidence_threshold": 0.80,
}
resp = client.post("/api/settings", json=update_payload)
assert resp.status_code == 200
# Verify updated settings
settings_resp = client.get("/api/settings")
assert settings_resp.status_code == 200
s_data = settings_resp.json()
assert s_data["downloads_dir"] == str(new_downloads)
assert s_data["movies_dir"] == str(new_movies)
assert s_data["dry_run"] is True
assert s_data["confidence_threshold"] == 0.80
# Verify test_env_file was written
assert test_env_file.is_file()
content = test_env_file.read_text(encoding="utf-8")
assert str(new_downloads) in content
assert str(new_movies) in content
def test_web_restart_endpoint(web_env, monkeypatch):
client, settings, downloads, movies, shows, test_env_file = web_env
# Prevent background task from exiting test process
monkeypatch.setattr("time.sleep", lambda _: None)
monkeypatch.setattr("os._exit", lambda _: None)
monkeypatch.setattr("os.execv", lambda *_: None)
resp = client.post("/api/restart")
assert resp.status_code == 200
assert resp.json()["status"] == "restarting"
def test_web_clear_batches_endpoint(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# Create sample files and run sort
client.post("/api/files/test-sample")
run_resp = client.post("/api/run", json={"dry_run": False})
assert run_resp.status_code == 200
batches = client.get("/api/batches").json()
assert len(batches) >= 1
# Clear batches
clear_resp = client.post("/api/batches/clear")
assert clear_resp.status_code == 200
assert clear_resp.json()["status"] == "cleared"
# Verify batches are now empty
batches_after = client.get("/api/batches").json()
assert len(batches_after) == 0
def test_folder_explorer_excludes_txt_files(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
from media_sorter.server import inspect_downloads_folder, list_files_in_dir
# Create real media files
(downloads / "Movie.2024.1080p.mkv").write_bytes(b"\x1aE\xdf\xa3" + b"\x00" * 100)
(shows / "Show.S01E01.mkv").write_bytes(b"\x1aE\xdf\xa3" + b"\x00" * 100)
# Create various .txt files in downloads and destination folders
(downloads / "Movie.2024.txt").write_text("info text")
(downloads / "readme.txt").write_text("readme text")
(downloads / "NOTES.TXT").write_text("notes text")
sub_dir = downloads / "Show.Release"
sub_dir.mkdir(parents=True, exist_ok=True)
(sub_dir / "tracker.txt").write_text("tracker info")
(shows / "show_notes.txt").write_text("notes")
# 1. inspect_downloads_folder must exclude all .txt files
inspected = inspect_downloads_folder(downloads, settings)
for f in inspected["files"]:
assert not f["name"].lower().endswith(".txt")
for s in inspected["shows"]:
for f in s["files"]:
assert not f["name"].lower().endswith(".txt")
for f in inspected["singles"]:
assert not f["name"].lower().endswith(".txt")
# 2. list_files_in_dir must exclude .txt files
shows_listed = list_files_in_dir(shows)
for f in shows_listed:
assert not f["name"].lower().endswith(".txt")
# 3. GET /api/files endpoint must exclude .txt files from folder explorer response
res = client.get("/api/files")
assert res.status_code == 200
data = res.json()
downloads_data = data["downloads"]
assert any(f["name"] == "Movie.2024.1080p.mkv" for f in downloads_data["files"])
assert not any(f["name"].lower().endswith(".txt") for f in downloads_data["files"])
assert not any(f["name"].lower().endswith(".txt") for f in downloads_data["singles"])
def test_folder_explorer_excludes_srt_files(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
from media_sorter.server import inspect_downloads_folder, list_files_in_dir
# Create real media files alongside .srt subtitles
(downloads / "Avatar.2009.1080p.mkv").write_bytes(b"\x1aE\xdf\xa3" + b"\x00" * 100)
(downloads / "Avatar.2009.1080p.en.srt").write_text("1\n00:00:01 --> 00:00:03\nSubtitles\n")
(downloads / "random_track.SRT").write_text("sub")
sub_dir = downloads / "Show.Episode.Folder"
sub_dir.mkdir(parents=True, exist_ok=True)
(sub_dir / "Episode 01.mkv").write_bytes(b"\x1aE\xdf\xa3" + b"\x00" * 100)
(sub_dir / "Episode 01.srt").write_text("sub")
(shows / "Show.S01E01.en.srt").write_text("sub")
# 1. inspect_downloads_folder must exclude all .srt files
inspected = inspect_downloads_folder(downloads, settings)
for f in inspected["files"]:
assert not f["name"].lower().endswith(".srt")
for s in inspected["shows"]:
for f in s["files"]:
assert not f["name"].lower().endswith(".srt")
for g in inspected["unsure_groups"]:
for f in g["files"]:
assert not f["name"].lower().endswith(".srt")
for f in inspected["singles"]:
assert not f["name"].lower().endswith(".srt")
# 2. list_files_in_dir must exclude .srt files
shows_listed = list_files_in_dir(shows)
for f in shows_listed:
assert not f["name"].lower().endswith(".srt")
# 3. GET /api/files endpoint must exclude .srt files from folder explorer response
res = client.get("/api/files")
assert res.status_code == 200
data = res.json()
downloads_data = data["downloads"]
assert any(f["name"] == "Avatar.2009.1080p.mkv" for f in downloads_data["files"])
assert not any(f["name"].lower().endswith(".srt") for f in downloads_data["files"])
assert not any(f["name"].lower().endswith(".srt") for f in downloads_data["singles"])
assert not any(f["name"].lower().endswith(".srt") for s in downloads_data["shows"] for f in s["files"])
assert not any(f["name"].lower().endswith(".srt") for g in downloads_data["unsure_groups"] for f in g["files"])
def test_delete_download_file_cleans_txt_and_parent_dir(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# Setup a subfolder in downloads with a file, a companion .txt, and an extra .txt
sub_dir = downloads / "Custom.Release.Folder"
sub_dir.mkdir(parents=True, exist_ok=True)
media = sub_dir / "sample.mkv"
media.write_text("sample")
companion_txt = sub_dir / "sample.txt"
companion_txt.write_text("companion")
extra_txt = sub_dir / "release.txt"
extra_txt.write_text("release note")
# Delete sample.mkv via API
del_res = client.delete(f"/api/files/download?name=Custom.Release.Folder/sample.mkv")
assert del_res.status_code == 200
assert del_res.json()["status"] == "deleted"
# Verify media, companion .txt, extra .txt, and empty parent subfolder are removed
assert not media.exists()
assert not companion_txt.exists()
assert not extra_txt.exists()
assert not sub_dir.exists()
assert downloads.exists()
def test_rollback_all_api(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# 1. Generate sample downloads and sort
client.post("/api/files/test-sample")
client.post("/api/run", json={"dry_run": False})
# Add extra file and sort again to form a second batch
f = downloads / "Extra.Film.2022.mkv"
f.write_text("extra movie")
client.post("/api/run", json={"dry_run": False})
# Call rollback all
rb_res = client.post("/api/rollback/all")
assert rb_res.status_code == 200
assert rb_res.json()["status"] == "ok"
assert rb_res.json()["reverted_files"] >= 4
# Verify files restored to downloads
files_res = client.get("/api/files").json()
assert len(files_res["downloads"]["files"]) >= 4
def test_manual_sort_file_api(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# Test sorting as movie with custom title & year
movie_file = downloads / "ambiguous_movie_file.mkv"
movie_file.write_text("movie payload")
res_movie = client.post("/api/files/manual-sort", json={
"relative_path": "ambiguous_movie_file.mkv",
"category": "movie",
"title": "Interstellar",
"year": 2014,
})
assert res_movie.status_code == 200
assert not movie_file.exists()
dest_movie = movies / "Interstellar (2014)" / "Interstellar (2014).mkv"
assert dest_movie.exists()
# Test sorting as TV show with custom title, season, episode
tv_file = downloads / "random_episode.mkv"
tv_file.write_text("tv payload")
res_tv = client.post("/api/files/manual-sort", json={
"relative_path": "random_episode.mkv",
"category": "tv",
"title": "Succession",
"season": 3,
"episode": 5,
})
assert res_tv.status_code == 200
assert not tv_file.exists()
dest_tv = shows / "Succession" / "Season 03" / "Succession - S03E05.mkv"
assert dest_tv.exists()
def test_quarantine_resolve_and_undo_api(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
from media_sorter.db import get_db_session
from media_sorter.models import QuarantineRecord, QuarantineStatus
from media_sorter.quarantine import QuarantineManager
quar_file = downloads / "unknown_sample.xyz"
quar_file.write_text("quarantine payload")
# Manually insert pending record
engine = init_db(db_path=settings.database.path)
with get_db_session(engine) as s:
qm = QuarantineManager(s)
rec = QuarantineRecord(
src=str(quar_file),
suggested_category="movie",
confidence=0.45,
reason="Unrecognized format",
status=QuarantineStatus.PENDING.value,
)
s.add(rec)
s.commit()
item_id = rec.id
# 1. GET /api/quarantine returns pending and resolved
q_res = client.get("/api/quarantine")
assert q_res.status_code == 200
q_data = q_res.json()
assert len(q_data["pending"]) == 1
assert q_data["pending"][0]["id"] == item_id
assert len(q_data["resolved"]) == 0
# 2. POST /api/quarantine/{id}/resolve with custom TV show title
res_resolve = client.post(f"/api/quarantine/{item_id}/resolve", json={
"category": "tv",
"title": "Severance",
"season": 1,
"episode": 1,
})
assert res_resolve.status_code == 200
assert not quar_file.exists()
dest_tv = shows / "Severance" / "Season 01" / "Severance - S01E01.xyz"
assert dest_tv.exists()
# Verify GET /api/quarantine reflects resolved item
q_res2 = client.get("/api/quarantine").json()
assert len(q_res2["pending"]) == 0
assert len(q_res2["resolved"]) == 1
# 3. POST /api/quarantine/{id}/undo restores file to src and marks PENDING
res_undo = client.post(f"/api/quarantine/{item_id}/undo")
assert res_undo.status_code == 200
assert res_undo.json()["status"] == "undone"
assert quar_file.exists()
assert not dest_tv.exists()
q_res3 = client.get("/api/quarantine").json()
assert len(q_res3["pending"]) == 1
assert len(q_res3["resolved"]) == 0
# 4. POST /api/quarantine/{id}/undo on pending item unflags/deletes it
res_unflag = client.post(f"/api/quarantine/{item_id}/undo")
assert res_unflag.status_code == 200
q_res4 = client.get("/api/quarantine").json()
assert len(q_res4["pending"]) == 0
def test_inspect_downloads_movie_subfolder_classified_as_movie(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# 1. Torrent-style movie inside a subfolder
movie_folder = downloads / "The.Dark.Knight.2008.1080p.BluRay.x264-ROVERS"
movie_folder.mkdir(parents=True, exist_ok=True)
movie_file = movie_folder / "The.Dark.Knight.2008.1080p.BluRay.x264-ROVERS.mkv"
movie_file.write_text("movie data")
# 2. Movie inside clean folder
opp_folder = downloads / "Oppenheimer (2023)"
opp_folder.mkdir(parents=True, exist_ok=True)
opp_file = opp_folder / "Oppenheimer.2023.2160p.mkv"
opp_file.write_text("oppenheimer data")
# 3. Legitimate TV show
tv_folder = downloads / "Breaking Bad Season 1"
tv_folder.mkdir(parents=True, exist_ok=True)
tv_file = tv_folder / "Breaking.Bad.S01E01.Pilot.mkv"
tv_file.write_text("tv show data")
# Query folder explorer
files_res = client.get("/api/files").json()
dl_info = files_res["downloads"]
# Movies must be in singles, not in shows
single_rel_paths = [s["relative_path"] for s in dl_info["singles"]]
assert "The.Dark.Knight.2008.1080p.BluRay.x264-ROVERS/The.Dark.Knight.2008.1080p.BluRay.x264-ROVERS.mkv" in single_rel_paths
assert "Oppenheimer (2023)/Oppenheimer.2023.2160p.mkv" in single_rel_paths
for s in dl_info["singles"]:
if "The.Dark.Knight" in s["relative_path"]:
assert s["detected_type"] == "movie"
assert "The Dark Knight" in s["believed_title"]
assert str(movies) in s["believed_destination"]
if "Oppenheimer" in s["relative_path"]:
assert s["detected_type"] == "movie"
assert "Oppenheimer" in s["believed_title"]
assert str(movies) in s["believed_destination"]
# Shows must only contain Breaking Bad, not the movies
show_names = [show["show_name"] for show in dl_info["shows"]]
assert "Breaking Bad" in show_names
assert "The Dark Knight" not in show_names
assert "The.Dark.Knight.2008.1080p.BluRay.x264-ROVERS" not in show_names
assert "Oppenheimer" not in show_names
def test_quarantine_bulk_resolve_and_bulk_undo_api(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
from media_sorter.db import get_db_session
from media_sorter.models import QuarantineRecord, QuarantineStatus
from media_sorter.quarantine import QuarantineManager
f1 = downloads / "Unknown.Movie.2021.1080p.mkv"
f1.write_text("movie1 payload")
f2 = downloads / "Another.Film.2023.720p.mkv"
f2.write_text("movie2 payload")
f3 = downloads / "Mystery.File.xyz"
f3.write_text("mystery payload")
engine = init_db(db_path=settings.database.path)
with get_db_session(engine) as s:
r1 = QuarantineRecord(src=str(f1), suggested_category="movie", confidence=0.4, reason="Low confidence", status=QuarantineStatus.PENDING.value)
r2 = QuarantineRecord(src=str(f2), suggested_category="movie", confidence=0.4, reason="Low confidence", status=QuarantineStatus.PENDING.value)
r3 = QuarantineRecord(src=str(f3), suggested_category="unknown", confidence=0.1, reason="Unrecognized format", status=QuarantineStatus.PENDING.value)
s.add_all([r1, r2, r3])
s.commit()
id1, id2, id3 = r1.id, r2.id, r3.id
# 1. Bulk resolve id1 and id2 as movie
res_b = client.post("/api/quarantine/bulk-resolve", json={
"item_ids": [id1, id2],
"category": "movie"
})
assert res_b.status_code == 200
assert res_b.json()["resolved_count"] == 2
assert not f1.exists()
assert not f2.exists()
assert f3.exists()
# 2. Bulk unflag id3
res_u = client.post("/api/quarantine/bulk-undo", json={
"item_ids": [id3],
"scope": "pending"
})
assert res_u.status_code == 200
assert res_u.json()["undone_count"] == 1
assert f3.exists() # Unflag leaves file in source
# Verify id3 is deleted from quarantine
q_data = client.get("/api/quarantine").json()
assert len(q_data["pending"]) == 0
assert len(q_data["resolved"]) == 2
# 3. Bulk undo all resolved items
res_all_undo = client.post("/api/quarantine/bulk-undo", json={
"scope": "resolved"
})
assert res_all_undo.status_code == 200
assert res_all_undo.json()["undone_count"] == 2
assert f1.exists()
assert f2.exists()
q_data_restored = client.get("/api/quarantine").json()
assert len(q_data_restored["pending"]) == 2
assert len(q_data_restored["resolved"]) == 0
def test_sort_show_endpoint_and_rollback(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# Create episodic show files and an unrelated file
ep1 = downloads / "Naruto Episode 001 Enter Naruto Uzumaki!.mkv"
ep2 = downloads / "Naruto Episode 002 My Name is Konohamaru!.mkv"
movie = downloads / "Inception (2010).mkv"
header = b"\x1aE\xdf\xa3" + b"\x00" * 300
ep1.write_bytes(header)
ep2.write_bytes(header)
movie.write_bytes(header)
# 1. Sort show specifically
res = client.post("/api/files/sort-show", json={
"show_name": "Naruto"
})
assert res.status_code == 200
data = res.json()
assert data["status"] == "ok"
assert data["moved_files"] == 2
assert data["show_name"] == "Naruto"
# Naruto episodes moved, movie remains in downloads
assert not ep1.exists()
assert not ep2.exists()
assert movie.exists()
# Destination directory contains organized show
naruto_dir = shows / "Naruto"
assert naruto_dir.exists()
organized_eps = list(naruto_dir.rglob("*.mkv"))
assert len(organized_eps) == 2
# 2. Rollback the show batch
batch_id = data["batch_id"]
rb_res = client.post("/api/rollback", json={"batch_id": batch_id})
assert rb_res.status_code == 200
assert rb_res.json()["reverted_files"] == 2
# Files restored to downloads
assert ep1.exists()
assert ep2.exists()
# 3. Non-existent show returns 404
err_res = client.post("/api/files/sort-show", json={
"show_name": "NonExistentShow"
})
assert err_res.status_code == 404
def test_dashboard_themes_and_rgb_feature(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
resp = client.get("/")
assert resp.status_code == 200
html = resp.text
# Verify all 30 themes are present in CSS
expected_themes = [
"cyber-dark", "oled-neon", "nord-frost", "dracula", "emerald-matrix",
"solar-sunset", "tokyo-night", "synthwave", "abyssal-ocean", "monokai-pro",
"terminal-crt", "paper-light", "neo-brutalism", "aurora-glass", "catppuccin-mocha",
"rose-pine", "gruvbox-dark", "solarized-dark", "nightowl", "vesper",
"bios-amber", "vapor-glitch", "mossy-stone", "crimson-eclipse", "blueprint-draft",
"neon-forest", "retro-retro", "golden-sand", "deep-space", "candy-cotton",
]
for theme_id in expected_themes:
assert f'[data-theme="{theme_id}"]' in html, f"Missing CSS for theme {theme_id}"
assert f'value="{theme_id}"' in html, f"Missing select option for theme {theme_id}"
assert f"id: '{theme_id}'" in html, f"Missing JS THEMES entry for theme {theme_id}"
# Verify RGB feature controls and styles
assert "rgb-mode" in html
assert "tab-rgb-mode-toggle" in html
assert "tab-rgb-speed-slider" in html
assert "toggleRGBMode" in html
assert "updateRGBSpeed" in html
assert "--rgb-duration" in html
def test_poster_local_storage_boundary_and_security(web_env, tmp_path: Path):
client, settings, downloads, movies, shows, test_env_file = web_env
# 1. Valid poster inside authorized storage root returns 200
valid_poster = shows / "poster.jpg"
valid_poster.write_bytes(b"\xff\xd8\xff\xe0" + b"image_data")
res_valid = client.get(f"/api/poster/local?path={valid_poster}")
assert res_valid.status_code == 200
# 2. Non-existent file inside storage root returns 404
missing_poster = shows / "missing.jpg"
res_missing = client.get(f"/api/poster/local?path={missing_poster}")
assert res_missing.status_code == 404
# 3. Path outside storage roots (traversal attempt) returns 403 Forbidden
outside_dir = tmp_path / "outside_secret"
outside_dir.mkdir()
secret_file = outside_dir / "secret.png"
secret_file.write_bytes(b"secret payload")
res_outside = client.get(f"/api/poster/local?path={secret_file}")
assert res_outside.status_code == 403
# 4. Non-image extension returns 400 Bad Request
text_file = downloads / "notes.txt"
text_file.write_text("not an image")
res_text = client.get(f"/api/poster/local?path={text_file}")
assert res_text.status_code == 400
def test_manual_sort_sanitization_and_reserved_names(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# 1. Path traversal in title is sanitized and contained
target_movie = downloads / "sample_movie.mkv"
target_movie.write_text("movie data")
res_traversal = client.post("/api/files/manual-sort", json={
"relative_path": "sample_movie.mkv",
"category": "movie",
"title": "../../../traversal_title",
"year": 2022,
})
assert res_traversal.status_code == 200
expected_dir = movies / "traversal_title (2022)"
assert expected_dir.exists()
assert (expected_dir / "traversal_title (2022).mkv").exists()
# 2. Windows reserved name in title (e.g. CON, AUX) is prefixed with underscore
target_tv = downloads / "con_show.mkv"
target_tv.write_text("tv data")
res_con = client.post("/api/files/manual-sort", json={
"relative_path": "con_show.mkv",
"category": "tv",
"title": "CON",
"season": 1,
"episode": 2,
})
assert res_con.status_code == 200
expected_tv_dir = shows / "_CON" / "Season 01"
assert expected_tv_dir.exists()
assert (expected_tv_dir / "_CON - S01E02.mkv").exists()
# 3. Pure forbidden characters title is rejected with 400
target_invalid = downloads / "invalid_file.mkv"
target_invalid.write_text("data")
res_bad = client.post("/api/files/manual-sort", json={
"relative_path": "invalid_file.mkv",
"category": "movie",
"title": ":::***???",
})
assert res_bad.status_code == 400
def test_process_locking_concurrency_409(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
from media_sorter.executor import acquire_process_lock
lock_file = settings.get_database_path().with_suffix(".lock")
with acquire_process_lock(lock_file):
# When lock is held, endpoints should return 409 Conflict
res_run = client.post("/api/run", json={"dry_run": True})
assert res_run.status_code == 409
res_rb = client.post("/api/rollback", json={"batch_id": "dummy"})
assert res_rb.status_code == 409
res_rba = client.post("/api/rollback/all")
assert res_rba.status_code == 409
dummy_file = downloads / "dummy.mkv"
dummy_file.write_text("test")
res_ms = client.post("/api/files/manual-sort", json={
"relative_path": "dummy.mkv",
"category": "movie",
"title": "Locked Movie",
})
assert res_ms.status_code == 409
def test_folder_explorer_correct_movie_and_show_buttons(web_env):
client, settings, downloads, movies, shows, test_env_file = web_env
# 1. HTML Dashboard contains correct movie & show buttons in Folder Explorer
resp = client.get("/")
assert resp.status_code == 200
assert "btn-correct-movie-explorer" in resp.text
assert "btn-correct-show-explorer" in resp.text
assert "bulk-btn-correct-movie" in resp.text
assert "bulk-btn-correct-show" in resp.text
assert "correctMovieByIndex" in resp.text
assert "correctShowByIndex" in resp.text
assert "correctMovieForShow" in resp.text
assert "correctMovieForGroup" in resp.text
assert "handleExplorerCorrectMovie" in resp.text
assert "handleExplorerCorrectShow" in resp.text
# 2. Test automatic/manual movie sort on a file in downloads
test_movie_file = downloads / "Unknown.Film.2023.1080p.mkv"
test_movie_file.write_bytes(b"dummy movie data")
sort_res = client.post("/api/files/manual-sort", json={
"relative_path": "Unknown.Film.2023.1080p.mkv",
"category": "movie",
"title": "Oppenheimer",
"year": 2023
})
assert sort_res.status_code == 200
expected_movie_file = movies / "Oppenheimer (2023)" / "Oppenheimer (2023).mkv"
assert expected_movie_file.exists()
# 3. Test automatic/manual show sort on a file in downloads
test_show_file = downloads / "Unknown.Show.S02E05.mkv"
test_show_file.write_bytes(b"dummy show data")
show_res = client.post("/api/files/manual-sort", json={
"relative_path": "Unknown.Show.S02E05.mkv",
"category": "tv",
"title": "Severance",
"season": 2,
"episode": 5
})
assert show_res.status_code == 200
expected_show_file = shows / "Severance" / "Season 02" / "Severance - S02E05.mkv"
assert expected_show_file.exists()
+164
View File
@@ -0,0 +1,164 @@
from pathlib import Path
import pytest
from media_sorter.tokenizer import FilenameTokenizer
@pytest.fixture
def tokenizer():
return FilenameTokenizer()
def test_tv_show_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("Breaking.Bad.S05E14.Ozymandias.1080p.BluRay.x264-ROVERS.mkv"))
assert tokens.is_episodic is True
assert tokens.title == "Breaking Bad"
assert tokens.season == 5
assert tokens.episode == 14
assert tokens.episode_title == "Ozymandias"
assert tokens.resolution == "1080p"
assert tokens.source == "BLURAY"
assert tokens.video_codec == "x264"
assert tokens.group == "ROVERS"
def test_tv_multi_episode(tokenizer):
tokens = tokenizer.tokenize(Path("Stranger.Things.S04E01-E02.Chapter.One.720p.WEB-DL.mkv"))
assert tokens.is_episodic is True
assert tokens.season == 4
assert tokens.episode == 1
assert tokens.multi_episodes == [1, 2]
assert tokens.resolution == "720p"
def test_anime_fansub_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("[SubsPlease] Frieren - Beyond Journey's End - 01 (1080p) [ABCD1234].mkv"))
assert tokens.is_anime is True
assert tokens.is_episodic is True
assert tokens.group == "SubsPlease"
assert "Frieren" in tokens.title
assert tokens.episode == 1
assert tokens.season == 1
def test_movie_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("Inception.2010.2160p.UHD.Remux.HEVC.TrueHD.Atmos-FraMeSToR.mkv"))
assert tokens.is_episodic is False
assert tokens.title == "Inception"
assert tokens.year == 2010
assert tokens.resolution == "2160p"
assert tokens.video_codec == "hevc"
assert tokens.audio_codec == "TRUEHD"
def test_music_track_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("/Music/Daft Punk - Discovery/02 - One More Time.flac"))
assert tokens.is_music is True
assert tokens.track == 2
assert tokens.title == "One More Time"
assert tokens.artist == "Daft Punk"
assert tokens.album == "Discovery"
def test_camera_and_date_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("IMG_20240815_142301.jpg"))
assert tokens.is_photo_or_home_video is True
assert tokens.date_stamp == "2024-08-15"
assert tokens.year == 2024
def test_podcast_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("The Daily - 2026-03-12 - The Sunday Read.mp3"))
assert tokens.artist == "The Daily"
assert tokens.year == 2026
assert tokens.date_stamp == "2026-03-12"
assert tokens.title == "The Sunday Read"
def test_movie_with_dimensions_not_episodic(tokenizer):
tokens = tokenizer.tokenize(Path("Interstellar.1920x1080.mkv"))
assert tokens.is_episodic is False
assert tokens.season is None
assert tokens.episode is None
assert tokens.resolution == "1080p"
assert "Interstellar" in tokens.title
tokens4k = tokenizer.tokenize(Path("Dune.Part.Two.3840x2160.mkv"))
assert tokens4k.is_episodic is False
assert tokens4k.season is None
assert tokens4k.episode is None
assert tokens4k.resolution == "2160p"
def test_movie_bracket_group_year_not_anime(tokenizer):
tokens = tokenizer.tokenize(Path("[YTS.MX] Movie Title - 2024 [1080p].mkv"))
assert tokens.is_anime is False
assert tokens.is_episodic is False
assert tokens.year == 2024
assert tokens.title == "Movie Title"
def test_standalone_episode_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("Naruto Episode 207 The Supposed Sealed Ability.mkv"))
assert tokens.is_episodic is True
assert tokens.title == "Naruto"
assert tokens.season == 1
assert tokens.episode == 207
assert tokens.episode_title == "The Supposed Sealed Ability"
def test_anime_fansub_without_group_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("BLEACH꞉ Sennen Kessen-hen - 27 E89717B7].mkv"))
assert tokens.is_anime is True
assert tokens.is_episodic is True
assert "BLEACH" in tokens.title
assert tokens.episode == 27
assert tokens.season == 1
def test_anime_ending_opening_tokenization(tokenizer):
tokens = tokenizer.tokenize(Path("[A&C] Sword Art Online Alicization S03ED01 [BDrip 1080p] [09F0EA6C].mkv"))
assert tokens.is_episodic is True
assert "Sword Art Online" in tokens.title
assert tokens.season == 3
assert tokens.episode == 1
def test_one_piece_tokenization_varieties(tokenizer):
# 1. Anime Time release with episode title and ancestor season folder
p1 = Path("One Piece (0001-1071+Movies+Specials)/Season 08 - Water Seven (229-263)/[Anime Time] One Piece - 0233 - Pirate Abduction Incident! A Pirate Ship That Can Only Await Its End! [1080p][HEVC 10bit x265][AAC].mkv")
t1 = tokenizer.tokenize(p1)
assert t1.title == "One Piece"
assert t1.season == 8
assert t1.episode == 233
assert "Pirate Abduction" in t1.episode_title
assert t1.group == "Anime Time"
# 2. Adjacent episode in same folder to ensure natural order
p2 = Path("One Piece (0001-1071+Movies+Specials)/Season 08 - Water Seven (229-263)/[Anime Time] One Piece - 0232 - Galley-La Company! A Grand Sight Dock 1!.mkv")
t2 = tokenizer.tokenize(p2)
assert t2.title == "One Piece"
assert t2.season == 8
assert t2.episode == 232
# 3. 4-digit episode S01E1171 WEB-DL
p3 = Path("One.Piece.S01E1171.1080p.CR.WEB-DL.AAC2.0.H.264.mkv")
t3 = tokenizer.tokenize(p3)
assert t3.title == "One Piece"
assert t3.season == 1
assert t3.episode == 1171
# 4. HatSubs release with 4-digit episode and CRC
p4 = Path("[HatSubs] One Piece 1073 (BD 1080p 10-bit Opus) v2 [194B3FBA].mkv")
t4 = tokenizer.tokenize(p4)
assert t4.title == "One Piece"
assert t4.episode == 1073
assert t4.crc32 == "194B3FBA"
# 5. Kaerizaki-Fansub underscore release
p5 = Path("[Kaerizaki-Fansub]_One_Piece_1076_[VOSTFR][FHD_1920x1080].mp4")
t5 = tokenizer.tokenize(p5)
assert t5.title == "One Piece"
assert t5.episode == 1076