Changelog
Every release of compress-media, newest first, read from the CHANGELOG in its repository — so what a version says here and what its code does cannot drift apart.
Added
- Make an animation (GIF maker). Turn 2–1000 still images (JPG, PNG, WebP, HEIC…) into one animated GIF, animated WebP or MP4.
- Web UI: a "Make an animation" mode with a frame strip. Reorder frames by dragging or with ← →, then pick the format, time per frame, loop count, size, fit (borders or crop), border colour and quality. Redo re-makes it with new settings.
- CLI:
compress-media animate [--format gif|webp|mp4] [--delay ms | --fps n] [--loop n] [--max-dim px] [--fit contain|cover] [--background hex] [--quality 1-100] [-o file]. Folders are sorted by name with numbers in order (shot-2beforeshot-10). - HTTP API:
POST /api/animations, taking multipart frames or chunked upload ids.GET /api/configreportsanimationMaxFrames.
- Subtitles. Turn the speech in a video or audio file into subtitles with times and text (SRT or WebVTT), ready to upload to YouTube, using whisper.cpp on your own server. Or put subtitles into the video.
- Speech recognition in about 100 languages, including Vietnamese: the language is detected automatically or chosen, and you can translate the subtitles into English. There are five speech models, from
tiny(75 MB) tolarge-v3-turbo(547 MB); each is downloaded once, on first use. Silence and music are skipped. - Subtitles are split into readable cues of at most two lines.
- Add subtitles to a video as a track viewers can turn on (no re-encode), or burn them into the picture for TikTok, Reels and other apps that ignore subtitle tracks.
- Use your own
.srt/.vttinstead of speech recognition. - Web UI: a "Subtitles" mode. Drop a video, or a video plus its
.srtwith the same name. View & edit shows the video with its subtitles next to an editor; Save makes the result again from your text. Redo switches between SRT, VTT, track and burn-in without listening again. - CLI:
compress-media subtitles [--lang vi] [--translate] [--model small] [--format srt|vtt] [--embed none|track|burn] [--font-size …] [--srt file]. - HTTP API:
POST /api/subtitlesandGET /api/jobs/:id/subtitles?format=srt|vtt. Jobs report astage(model,transcribe,embed). - Configuration:
WHISPER_MODEL,WHISPER_MODELS_DIR,WHISPER_DOWNLOAD(off for servers without internet access),WHISPER_MODEL_URL,WHISPER_PATH,WHISPER_THREADS.
- Speech recognition in about 100 languages, including Vietnamese: the language is detected automatically or chosen, and you can translate the subtitles into English. There are five speech models, from
Changed
- The mode switch (Compress files, Make an animation, Subtitles) sits at the top of the page, above the settings and the file list.
- The Docker image includes whisper.cpp and the DejaVu fonts (about 100 MB more;
--build-arg WHISPER=falseleaves whisper.cpp out). Speech models are kept in/data/models, on the data volume.
Fixed
- Animated GIF or WebP input is no longer reduced to its first frame when the chosen image format can't animate. AVIF output becomes animated WebP, and JPEG/PNG output stays GIF. The result says so in the UI, the CLI report and
info.note.
Breaking
- Login is required by default. Set
AUTH_USERNAME(a username or email) andAUTH_PASSWORD, or turn login off withAUTH_ENABLED=false. Without a password, one is generated and kept across restarts. Scripts sign in with HTTP Basic or a bearer token (AUTH_TOKEN).
Added
- Built-in login: signed session cookies, HTTP Basic, bearer token, and a rate limit on failed attempts.
- Video:
- trim to a start and end time;
- WebM (VP9 or AV1 + Opus) and AV1 in MP4;
- video to animated GIF;
- two-pass encoding for a target size in MB;
- GPU encoders detected at startup: NVIDIA NVENC, Intel Quick Sync, VA-API and AMD AMF, next to Apple VideoToolbox.
- PDF compression with Ghostscript, with presets from "smallest" to "prepress". Ghostscript is included in the default image;
-nopdfimages leave it out. - Compress in the browser. An optional mode using WebCodecs (Mediabunny), where files never leave the device. Anything the browser can't do falls back to the server, with a note.
- Download all results as one ZIP.
- Integrations:
- live progress over Server-Sent Events;
- webhooks signed with HMAC, with retries and SSRF protection;
- batch status for many jobs in one call.
- Scale out: an optional Redis/BullMQ queue (
QUEUE=redis),ROLE=web|worker, and acompress-media workercommand. - With object storage, inputs stay in the bucket, so workers on other hosts can fetch them.
- JSDoc types checked with
tsc(npm run typecheck). THIRD_PARTY_NOTICES.md, and OCI license labels on the images (FFmpeg GPL, Ghostscript AGPL, libvips LGPL, Mediabunny MPL).
Fixed
- AVIF input was detected as HEIF.
Added
- Chunked, resumable uploads, modelled on S3 multipart uploads, for files of several GB. The part size (
UPLOAD_PART_MB) and parallelism (UPLOAD_CONCURRENCY) are configurable, and interrupted uploads resume after a reload. - Object storage:
STORAGE=s3works with any S3-compatible bucket (AWS S3, Cloudflare R2, Backblaze B2, DigitalOcean Spaces, SeaweedFS…). Browsers upload straight to the bucket with presigned URLs. There is a CORS check at startup, andS3_SETUP_CORScan set it up.
First public release.
- Web UI in English and Vietnamese: drag and drop, presets, before/after preview, redo with new settings, cancel, and download all.
- Video (MOV, MP4, MKV…) to H.264/H.265 MP4; images (JPG, PNG, HEIC, WebP, AVIF, GIF) to JPEG, WebP, AVIF or PNG; audio to MP3, M4A or Opus.
- CLI:
compress,probe,infoandserve, with a JSON report and folder mirroring. - HTTP API for scripts and integrations.
- Docker image for
linux/amd64andlinux/arm64, and Docker Compose files. - Docs for people (README,
docs/) and for AI agents (skills/,AGENTS.md,llms.txt). - Tests: node:test for the API and CLI, and Playwright end-to-end tests on Chromium, Firefox and WebKit.