navidrome/core/maintenance.go
ts c6732e1fdf
feat(cli): add missing file list and remap subcommands (#5928)
* feat(cli): add missing file list and remap subcommands

Signed-off-by: zerovox <933064+zerovox@users.noreply.github.com>

* fix: prevent remapping from dropping participants on target track

* fix: after remapping, refresh stats synchronously

* fix: only move album annotations if moving a track would empty the old album

* fix(persistence): keep the new item's annotation when reassigning onto an item the user already annotated

ReassignAnnotation was a plain UPDATE; the annotation table is unique on
(user_id, item_id, item_type), so when a user had annotated both items the
statement aborted and none of the rows moved. In the scanner that surfaced as
a warning; in the missing-file remap it rolled back the whole operation.
UPDATE OR IGNORE moves what it can and leaves the conflicting rows for GC.

* fix(core): keep the target track's history when remapping a missing file onto it

The remap discards the target's row, and GC then dropped its play counts,
stars, ratings, bookmarks and every playlist entry pointing at it. That is
harmless in the scanner, whose target was imported seconds earlier, but the
CLI lets the user pick any existing track. Move those references onto the
surviving id first; where a user already has a row for both, theirs on the
missing file wins.

* fix(persistence): stop FindByPaths dropping plain paths that contain a colon

Any colon was taken as the libraryID separator, and a non-numeric prefix
made the whole path vanish from the lookup. 'missing fix' then rejected the
very paths 'missing list' printed, and M3U imports silently skipped such
tracks. Only a numeric prefix qualifies a path now.

* perf(cli): stream 'missing list' instead of loading every missing file into memory

GetAll materialised the whole result set before a single row was written;
on a library with 97k missing files that peaked at 1.28 GB of RSS. Iterate
the repository cursor and write rows as they arrive.

* refactor(core): tidy the missing-file remap

Drop the log lines copied from deleteMissing that still said 'after deleting
missing files', the debug-on-success branches, and the what-comments; build
the affected album list without slice helpers.

* fix(cli): move path to the last column of 'missing list'

Path is the only variable-width field, so leading with it misaligns every
row that follows. Applies to both csv and json.

* fix(persistence): also try a numeric colon prefix as a plain path

'1999: A Different Life/01.mp3' parsed as library 1999 plus a truncated path
and matched nothing. The prefix is ambiguous, so search both ways.

Also buffer the json branch of 'missing list', which wrote a syscall per row.

* fix(persistence): move scrobbles and buffered scrobbles off a discarded media file

Both tables carry ON DELETE CASCADE on media_file_id, so 'missing fix'
deleting the target erased its play history and dropped scrobbles still
waiting on an external service. scrobble_buffer needs OR IGNORE for its
unique (user_id, service, media_file_id, play_time).

* fix(persistence): recompute the cached average rating after merging annotations

Merging the discarded row's annotations grows the rating population of the
surviving track, so media_file.average_rating no longer matched what the
annotation rows say. Only reachable since the remap started merging those
rows instead of deleting them.

* fix(persistence): recompute the cached average rating inside ReassignAnnotation

Moving annotation rows always changes the new item's rating population, so
the recompute belongs with the move rather than at each call site. Covers
the album reassign in the remap and the two scanner sites, and replaces the
explicit call ReassignReferences was making.

Album was the worse case: rate an album, move its files, and 'missing fix'
handed the rating to an album still caching an average of 0.

* fix(cli): let libraryID:path win over a file literally named like one

FindByPaths searches a numeric-prefixed reference both ways, so a top-level
file named '1:foo.mp3' can tie with library 1's 'foo.mp3'. The CLI then
rejected the reference as ambiguous while advising the exact syntax the
caller had used. Also disambiguates the same path in two libraries, which
is what the qualified form is for.

---------

Signed-off-by: zerovox <933064+zerovox@users.noreply.github.com>
Co-authored-by: Deluan Quintão <deluan@navidrome.org>
2026-09-12 12:08:25 -04:00

319 lines
10 KiB
Go

package core
import (
"context"
"errors"
"fmt"
"slices"
"sync"
"time"
"github.com/Masterminds/squirrel"
"github.com/navidrome/navidrome/log"
"github.com/navidrome/navidrome/model"
"github.com/navidrome/navidrome/model/request"
"github.com/navidrome/navidrome/utils/slice"
)
var (
// ErrNotMissing is returned when a remap is attempted from a file not marked as missing.
ErrNotMissing = errors.New("file is not marked as missing")
// ErrTargetMissing is returned when the remap target is itself a missing file.
ErrTargetMissing = errors.New("target file is missing")
// ErrSameFile is returned when the remap source and target are the same file.
ErrSameFile = errors.New("missing and target are the same file")
)
type Maintenance interface {
// DeleteMissingFiles deletes specific missing files by their IDs
DeleteMissingFiles(ctx context.Context, ids []string) error
// DeleteAllMissingFiles deletes all files marked as missing
DeleteAllMissingFiles(ctx context.Context) error
// RemapMissingFile moves a missing file's identity onto an existing file, the manual
// counterpart to the scanner's move detection (phaseMissingTracks.moveMatched).
RemapMissingFile(ctx context.Context, missingID, targetID string) error
}
type maintenanceService struct {
ds model.DataStore
wg sync.WaitGroup
}
func NewMaintenance(ds model.DataStore) Maintenance {
return &maintenanceService{
ds: ds,
}
}
func (s *maintenanceService) DeleteMissingFiles(ctx context.Context, ids []string) error {
return s.deleteMissing(ctx, ids)
}
func (s *maintenanceService) DeleteAllMissingFiles(ctx context.Context) error {
return s.deleteMissing(ctx, nil)
}
func (s *maintenanceService) RemapMissingFile(ctx context.Context, missingID, targetID string) error {
if missingID == targetID {
return fmt.Errorf("%w: %q", ErrSameFile, missingID)
}
missing, err := s.ds.MediaFile(ctx).Get(missingID)
if err != nil {
return fmt.Errorf("loading missing file %q: %w", missingID, err)
}
if !missing.Missing {
return fmt.Errorf("%w: %q", ErrNotMissing, missingID)
}
target, err := s.ds.MediaFile(ctx).GetWithParticipants(targetID)
if err != nil {
return fmt.Errorf("loading target file %q: %w", targetID, err)
}
if target.Missing {
return fmt.Errorf("%w: %q", ErrTargetMissing, targetID)
}
oldAlbumID, newAlbumID := missing.AlbumID, target.AlbumID
err = s.ds.WithTx(func(tx model.DataStore) error {
discardedID := target.ID
// Preserve the original created_at so the remapped track doesn't resurface in "Recently Added"
target.CreatedAt = missing.CreatedAt
target.ID = missing.ID
if err := tx.MediaFile(ctx).Put(target); err != nil {
return fmt.Errorf("update matched track: %w", err)
}
// Unlike the scanner's freshly-imported target, this one may carry history of its own
if err := tx.MediaFile(ctx).ReassignReferences(discardedID, missing.ID); err != nil {
return fmt.Errorf("reassign target references: %w", err)
}
if err := tx.MediaFile(ctx).Delete(discardedID); err != nil {
return fmt.Errorf("delete discarded track: %w", err)
}
if oldAlbumID != newAlbumID {
oldAlbumTracks, err := tx.MediaFile(ctx).CountAll(model.QueryOptions{Filters: squirrel.Eq{"album_id": oldAlbumID}})
if err != nil {
return fmt.Errorf("get old album tracks: %w", err)
}
if oldAlbumTracks == 0 {
if err := tx.Album(ctx).ReassignAnnotation(oldAlbumID, newAlbumID); err != nil {
return fmt.Errorf("reassign album annotations: %w", err)
}
if err := tx.Album(ctx).CopyAttributes(oldAlbumID, newAlbumID, "created_at"); err != nil && !errors.Is(err, model.ErrNotFound) {
return fmt.Errorf("copy album attributes: %w", err)
}
}
}
return nil
})
if err != nil {
log.Error(ctx, "Error remapping missing file", "missing", missing.Path, "target", target.Path, err)
return err
}
if err := s.ds.GC(ctx); err != nil {
log.Error(ctx, "Error running GC after remapping missing file", err)
return err
}
// Stats are refreshed synchronously, unlike deleteMissing, so the CLI sees them before it exits.
// album/artist play count aggregates are not recalculated here; they are refreshed by the next scan.
if _, err := s.ds.Artist(ctx).RefreshStats(true); err != nil {
log.Error(ctx, "Error refreshing artist stats after remapping missing file", err)
}
affectedAlbumIDs := []string{newAlbumID}
if oldAlbumID != newAlbumID {
affectedAlbumIDs = append(affectedAlbumIDs, oldAlbumID)
}
if err := s.refreshAlbums(ctx, affectedAlbumIDs); err != nil {
log.Error(ctx, "Error refreshing album stats after remapping missing file", err)
}
return nil
}
// deleteMissing handles the deletion of missing files and triggers necessary cleanup operations
func (s *maintenanceService) deleteMissing(ctx context.Context, ids []string) error {
// Track affected album IDs before deletion for refresh
affectedAlbumIDs, err := s.getAffectedAlbumIDs(ctx, ids)
if err != nil {
log.Warn(ctx, "Error tracking affected albums for refresh", err)
// Don't fail the operation, just log the warning
}
// Delete missing files within a transaction
err = s.ds.WithTx(func(tx model.DataStore) error {
if len(ids) == 0 {
_, err := tx.MediaFile(ctx).DeleteAllMissing()
return err
}
return tx.MediaFile(ctx).DeleteMissing(ids)
})
if err != nil {
log.Error(ctx, "Error deleting missing tracks from DB", "ids", ids, err)
return err
}
// Run garbage collection to clean up orphaned records
if err := s.ds.GC(ctx); err != nil {
log.Error(ctx, "Error running GC after deleting missing tracks", err)
return err
}
// Refresh statistics in background. album/artist play count aggregates are not recalculated
// here; they are refreshed by the next scan.
s.refreshStatsAsync(ctx, affectedAlbumIDs)
return nil
}
// refreshAlbums recalculates album attributes (size, duration, song count, etc.) from media files.
// It uses batch queries to minimize database round-trips for efficiency.
func (s *maintenanceService) refreshAlbums(ctx context.Context, albumIDs []string) error {
if len(albumIDs) == 0 {
return nil
}
log.Debug(ctx, "Refreshing albums", "count", len(albumIDs))
// Process in chunks to avoid query size limits
const chunkSize = 100
for chunk := range slice.CollectChunks(slices.Values(albumIDs), chunkSize) {
if err := s.refreshAlbumChunk(ctx, chunk); err != nil {
return fmt.Errorf("refreshing album chunk: %w", err)
}
}
log.Debug(ctx, "Successfully refreshed albums", "count", len(albumIDs))
return nil
}
// refreshAlbumChunk processes a single chunk of album IDs
func (s *maintenanceService) refreshAlbumChunk(ctx context.Context, albumIDs []string) error {
albumRepo := s.ds.Album(ctx)
mfRepo := s.ds.MediaFile(ctx)
// Batch load existing albums
albums, err := albumRepo.GetAll(model.QueryOptions{
Filters: squirrel.Eq{"album.id": albumIDs},
})
if err != nil {
return fmt.Errorf("loading albums: %w", err)
}
// Create a map for quick lookup
albumMap := make(map[string]*model.Album, len(albums))
for i := range albums {
albumMap[albums[i].ID] = &albums[i]
}
// Batch load all media files for these albums
mediaFiles, err := mfRepo.GetAll(model.QueryOptions{
Filters: squirrel.Eq{"album_id": albumIDs},
Sort: "album_id, path",
})
if err != nil {
return fmt.Errorf("loading media files: %w", err)
}
// Group media files by album ID
filesByAlbum := make(map[string]model.MediaFiles)
for i := range mediaFiles {
albumID := mediaFiles[i].AlbumID
filesByAlbum[albumID] = append(filesByAlbum[albumID], mediaFiles[i])
}
// Recalculate each album from its media files
for albumID, oldAlbum := range albumMap {
mfs, hasTracks := filesByAlbum[albumID]
if !hasTracks {
// Album has no tracks anymore, skip (will be cleaned up by GC)
log.Debug(ctx, "Skipping album with no tracks", "albumID", albumID)
continue
}
// Recalculate album from media files
newAlbum := mfs.ToAlbum()
// Only update if something changed (avoid unnecessary writes)
if !oldAlbum.Equals(newAlbum) {
// Preserve original timestamps
newAlbum.UpdatedAt = time.Now()
newAlbum.CreatedAt = oldAlbum.CreatedAt
if err := albumRepo.Put(&newAlbum); err != nil {
log.Error(ctx, "Error updating album during refresh", "albumID", albumID, err)
// Continue with other albums instead of failing entirely
continue
}
log.Trace(ctx, "Refreshed album", "albumID", albumID, "name", newAlbum.Name)
}
}
return nil
}
// getAffectedAlbumIDs returns distinct album IDs from missing media files
func (s *maintenanceService) getAffectedAlbumIDs(ctx context.Context, ids []string) ([]string, error) {
var filters squirrel.Sqlizer = squirrel.Eq{"missing": true}
if len(ids) > 0 {
filters = squirrel.And{
squirrel.Eq{"missing": true},
squirrel.Eq{"media_file.id": ids},
}
}
mfs, err := s.ds.MediaFile(ctx).GetAll(model.QueryOptions{
Filters: filters,
})
if err != nil {
return nil, err
}
// Extract unique album IDs
albumIDMap := make(map[string]struct{}, len(mfs))
for _, mf := range mfs {
if mf.AlbumID != "" {
albumIDMap[mf.AlbumID] = struct{}{}
}
}
albumIDs := make([]string, 0, len(albumIDMap))
for id := range albumIDMap {
albumIDs = append(albumIDs, id)
}
return albumIDs, nil
}
// refreshStatsAsync refreshes artist and album statistics in background goroutines
func (s *maintenanceService) refreshStatsAsync(ctx context.Context, affectedAlbumIDs []string) {
// Refresh artist stats in background
s.wg.Go(func() {
bgCtx := request.AddValues(context.Background(), ctx)
if _, err := s.ds.Artist(bgCtx).RefreshStats(true); err != nil {
log.Error(bgCtx, "Error refreshing artist stats after deleting missing files", err)
} else {
log.Debug(bgCtx, "Successfully refreshed artist stats after deleting missing files")
}
// Refresh album stats in background if we have affected albums
if len(affectedAlbumIDs) > 0 {
if err := s.refreshAlbums(bgCtx, affectedAlbumIDs); err != nil {
log.Error(bgCtx, "Error refreshing album stats after deleting missing files", err)
} else {
log.Debug(bgCtx, "Successfully refreshed album stats after deleting missing files", "count", len(affectedAlbumIDs))
}
}
})
}
// Wait waits for all background goroutines to complete.
// WARNING: This method is ONLY for testing. Never call this in production code.
// Calling Wait() in production will block until ALL background operations complete
// and may cause race conditions with new operations starting.
func (s *maintenanceService) wait() {
s.wg.Wait()
}