Rembg
Image ProcessingPortrait Background Removal
Rembg removes image backgrounds automatically. Use it as a CLI, Python library, HTTP server, or Docker container.
Pre-configured local AI bundles — cloud drive downloads, scan to save on mobile, ready out of the box.
34 tools · Google Drive & Mega · Browse categories
Showing 34 of 34 tools
Portrait Background Removal
Rembg removes image backgrounds automatically. Use it as a CLI, Python library, HTTP server, or Docker container.
Document OCR & Layout
Converts documents and images into structured, AI-friendly data such as JSON and Markdown.
Real-time Digital Human
Real-time talking-head generation from audio and a single portrait image, developed by Soul AI Lab.
Text-to-Speech
Open-source TTS from Alibaba Qwen team. Supports expressive streaming speech, voice design, and voice cloning.
Lip Sync for Digital Humans
Generate lip-synced talking-head videos from audio and reference portraits.
Text-to-Video
Generate short videos from text prompts with the LTX 2.3 text-to-video pipeline.
Inpainting & Watermark Removal
Remove watermarks, fill missing regions, and erase objects from images with inpainting models.
Motion Transfer
Transfer motion from a reference video to a still image and generate animated results.
Vision-Language OCR
Accurate layout parsing for skewed, scanned, low-light, and phone-captured documents.
Image/PDF to Markdown
Structured document parsing framework that turns images and PDFs into pixel-accurate Markdown output.
Virtual Try-On
One-click virtual outfit try-on for fashion and e-commerce workflows.
Video Background Removal
Remove backgrounds from videos with high-quality matting for compositing and editing.
Multi-Angle Face Edit
Edit faces from multiple angles with Qwen image editing models.
Face Swap
Industry-popular face swapping toolkit with CUDA acceleration for local GPU inference.
Head Swap
Lightweight head-swap workflow built on FLUX.2 Klein for fast local generation.
Speech to Subtitles
Fast Whisper-based transcription for turning speech into timed subtitles locally.
Aligned Transcription
Whisper transcription with word-level timestamps and speaker diarization support.
Speech Recognition
Automatic speech recognition pack for converting audio into text and subtitles.
Garment Extraction
Extract clothing items from photos for virtual try-on and catalog workflows.
Image Segmentation
Segment Anything Model for interactive and automatic object segmentation in images.
Streaming Transcription
Dockerized streaming ASR for real-time speech-to-text on local hardware.
Open-source Digital Human
Open-source AI avatar pipeline for creating talking digital humans locally.
Edge Speech Toolkit
All-in-one edge speech stack: ASR, TTS, VAD, and speaker ID with ONNX runtime.
Offline Speech Synthesis
Fast, lightweight offline text-to-speech engine with many voice models.
Offline Speech Recognition
Open-source offline speech recognition toolkit for 20+ languages on CPU.
Lightweight Offline TTS
Compact ONNX-based TTS model for low-resource offline speech synthesis.
CPU-friendly OCR
Lightweight OCR engine that runs well on CPU without a dedicated GPU.
PDF to Markdown
High-quality PDF and document conversion to clean Markdown with layout preservation.
Image & Video Upscale
Classic open-source super-resolution for enhancing image and video quality.
Face Restoration
Restore and enhance faces in old or low-quality photos and videos.
Photo Colorization
Colorize black-and-white photos and footage with deep learning models.
Classic OCR Engine
The classic open-source OCR engine supporting 100+ languages offline.
Lightweight Image Generation
Compact 4B-parameter FLUX.2 Klein model for fast local AI image generation.
PDF Parsing
Open-source PDF parsing toolkit for extracting structured content from documents.
Free ready-to-run local AI tool packs — background removal, OCR, speech-to-text, digital humans, face swap, and more. Google Drive & Mega downloads.