compress-media
A self-hosted compressor for video, images and audio. Drag files onto a page, run one command, or call it from your own code — on a machine you own, with nothing uploaded to anybody else.
docker run -d --name compress-media -p 127.0.0.1:4747:4747 \
-v compress-media-data:/data runsnip/compress-mediaThen open localhost:4747. Node 20 and up works too, without Docker.
From 2.0 that page asks you to sign in. Set AUTH_USERNAME and AUTH_PASSWORD, or pass AUTH_ENABLED=false to keep it open on a machine only you can reach.
What it takes, and what it gives back
Video
- From
- MOV, MP4, MKV, WebM, AVI
- To
- MP4, in H.264 or H.265
A quality preset, or a target size in megabytes when something has to fit an upload limit.
Images
- From
- JPG, PNG, WebP, AVIF, HEIC, GIF, TIFF
- To
- JPEG, WebP, AVIF or PNG
Resize by longest edge and convert format in one pass. HEIC from a phone comes out as something a browser will open.
Audio
- From
- WAV, M4A, FLAC, AIFF, MP3
- To
- MP3, M4A or Opus
Set a bitrate, or fold a stereo recording of one voice down to mono.
- From
- To
Presets from smallest to print quality, through Ghostscript. The -nopdf images leave it out.
Animations
- From
- 2 to 1000 stills — JPG, PNG, WebP, HEIC
- To
- Animated GIF, animated WebP or MP4
Reorder the frames, set the time each is held and how often it loops, then pick a size and how to fit it.
Subtitles
- From
- Any video or audio with speech
- To
- SRT or WebVTT, or burnt into the picture
Speech recognition on your own server with whisper.cpp, in about a hundred languages, Vietnamese among them.
Four ways in, one engine behind them
ffmpeg for video and audio, sharp for images. Everything below is a different door onto the same job queue.
A web page you drop files on
Drag a batch in or paste it. Uploads go in 8 MB parts, four at a time, and survive a dropped connection. Progress and an estimate while it works, then the result beside the original to compare before you keep it.
A command, for the other times
compress-media <file> with the flags you would expect — --max-res, --target-mb, --codec, --bitrate. compress-media probe <file> reads what a file already is. No server needs to be running.
An HTTP API, for your own software
Chunked upload, create a job, poll it, download the result. The same endpoints the web page uses, because it is the same server.
An agent, if you would rather ask
It ships skills for Claude and other assistants: compress these files and never touch the originals, or deploy the service and put a reverse proxy in front of it.
It is not on the internet unless you put it there
There is no sign-in, and that is exactly why it binds to 127.0.0.1 and stays there until you decide otherwise. A file’s name is never used as a path, and the server deletes nothing outside its own uploads and outputs folders.
Or give it a bucket and stand aside
With STORAGE=s3 the browser uploads straight to your own bucket through presigned URLs, and the server never holds the bytes. S3, R2, Google Cloud Storage, MinIO, SeaweedFS, Backblaze B2, DigitalOcean Spaces and Wasabi.
What the 2.x releases added
- Built-in login, on by default
- A username and password through AUTH_USERNAME and AUTH_PASSWORD, or AUTH_ENABLED=false to keep it open. Scripts sign in with HTTP Basic or an API token.
- More video out
- Trim to a start and end time. AV1 in MP4, WebM with VP9 or AV1 and Opus, and an animated GIF from any video.
- Size targeting that lands closer
- Two passes when you ask for a target in megabytes, so the result sits nearer the limit instead of under it.
- GPU encoding beyond macOS
- NVIDIA NVENC, Intel Quick Sync, VA-API and AMD AMF on Linux and Windows, found automatically. VideoToolbox stays.
- PDFs too
- Presets from smallest to print quality, through Ghostscript.
- Compressing without uploading
- An optional mode where video, audio and images are compressed on the visitor's own device with WebCodecs — nothing leaves it at all.
- Everything in one ZIP
- Download a whole batch as a single archive rather than file by file.
- Things to build against
- Live progress over Server-Sent Events, webhooks signed with HMAC when a job finishes, and the status of many jobs in one call.
- More than one machine
- An optional Redis and BullMQ queue, so several web servers and separate workers share the load.
- Speech models, kept on your server
- Five whisper.cpp models from tiny at 75 MB to large-v3-turbo at 547 MB, each downloaded once on first use. WHISPER_DOWNLOAD=false for a server with no way out to the internet.
- Animated input is no longer flattened
- An animated GIF or WebP asked to become a format that cannot animate now becomes animated WebP, or stays GIF, instead of quietly losing every frame but the first.
Upgrading from 1.x: 2.0 turns login on by default. Scripts that call it today will need credentials — or AUTH_ENABLED=false — after the update. That is what makes it 2.0 rather than 1.2.
Every release, and what changed in eachPlanned for 2.2
Next up; not built yet, and details may change.
- Cleaner sound
- Loudness normalisation to the level YouTube and TikTok expect (−14 LUFS), and background-noise reduction — fans, air conditioning — for video and audio.
- Rotate, flip and change speed
- Turn a sideways video upright, mirror it, or play it 1.5×, 2× or as a timelapse.
- Cut the silences
- Remove long pauses from talking videos automatically, using the same speech detection as the subtitles.
- Watermark
- Put a logo or text on videos and images, with a position and an opacity.
- YouTube chapters
- Chapter timestamps suggested from the subtitles, ready to paste into the description.
- Subtitles in other languages
- Translate them into more than English — English to Vietnamese, for instance.
- Join videos
- Put several clips together into one.
Same idea, smaller scale
compress-media is a tool you run yourself instead of paying per file. RunSnip is the other half of that habit: somewhere to build and publish a small web project without installing anything, and an MCP server so your agent can build one for you. Both are made by the same people, and both are meant to be understood rather than trusted.
The code is MIT. The builds of ffmpeg, x264 and x265 it uses carry the GPL, and sharp and libvips carry Apache-2.0 and LGPL-3.0 — worth knowing before it goes anywhere near a product you sell.