Krea 2 Identity Edit
Identity-preserving instruction image editing on Krea 2
Apps started here then claimed by the 🐐 authors
Identity-preserving instruction image editing on Krea 2
Efficient native-resolution image generation and editing
Codec-native video & image understanding with Mage-VL 4B
Extend images into larger canvases with Krea 2 outpaint
Live interactive world rollout from an image
Multi-view character sheet from one image (FLUX.2 LoRA)
Distilled LTX-2.3 identity video from a reference photo
Unified audio-text intelligence
Realtime VLM for image and video understanding
Multilingual CPU-only ASR with a 1.58-bit BitNet decoder
Bilingual EN/KO speech LM - transcribe, ask, and speak
Infographic generation & editing with SenseNova-U1 V3
Verbatim + intended transcripts with word-level timing
Zero-shot TTS with explicit word-level prosody control
Aerial object detection - YOLO Models
Turn a 3D mesh into a parametric CadQuery program
Generative 4x video upscaling with LTX-2.3 IC-LoRA
Depth-controlled image generation with Krea-2 Turbo
Re-render a video from a new camera angle via IC-LoRA
Remove people & vehicles from video, keep the background
1.06M-param pure-loop transformer with six effort levels
Style-guided image generation with Krea 2 Turbo
FLUX.2 Klein 9B fine-tune for image generation and editing
Parse document images into structured Markdown
Krea 2 Turbo HD text-to-image with enhanced VAE
Native size-sensitive matroid audit
Hierarchical parallel document parsing with a 1B VLM
Edit images or generate from Canny edges with NK2E on Krea 2
DINOv3 feature upsampling with ViT-Up
Tiny any-to-any multimodal GPT (text + image) prototype demo
LightOn multimodal document reranker with pointwise scoring
Cinematic product commercial style video LoRA for LTX-2.3
Segment and track object instances in video with QueenVIS
Occult philosophy chat LLM fine-tuned on Gemma 4 12B
Multilingual PII detection & LLM safety moderation
Tiny 31.7M SLM that solves arithmetic expressions
Generate scene-linear HDR images with Krea 2 + LogC4 LoRA
ClinFusion medical multimodal LLM for 2D images
Speech-driven talking-head video (Bernini-R S2V, Wan2.2)
Four voices in one 10M-param CPU model, with blending
Listwise document reranker with jina-reranker-v3.5
Photorealistic skin-texture LoRA for Krea 2 Turbo
MobileWan text-to-video generation (Qualcomm AI Research)
Subject-driven text-to-video from reference images (Wan2.2)
Arabic speech to fully-diacritized text (tashkeel)
SAGE retrieves matching place images from a gallery
Convert molecular structure images into E-SMILES strings
AlienLM privacy layer for black-box LLMs
Compact 60M T5 translator for 15 languages
Unified AR model for image understanding & generation
Faithful x4 image super-resolution via FLUX.1-dev dual-LoRA
Zero-shot multilingual NER with GLiNER on LFM2.5-350M
Multi-speaker meeting transcription with diarization
Compact multilingual ASR model (324M params)
Speak to a drone using an image, video, text, or voice.
Image and video understanding with MOSS-VL multimodal model
Medical reasoning VLM based on Qwen3.6-27B
Recognize text from WordArt / artistic scene text images
Agentic video understanding with active perception loop
Multimodal bio foundation model for molecules & proteins
Forecast 3D point trajectories from video and language
VLA model for scientific laboratory robotics
Distractor-free 2D enhancer for radiance field renderings
Fill-mask demo for a French medical ModernBERT
Speech-driven lip-sync on Wan 2.2 image-to-video
3D-centric world-spatial-action robot policy demo
Multi-shot narrated video with cross-shot memory
Multi-image instruction-guided image editing
Encode a context into a LoRA and chat with Qwen3-8B
Turn hardware ideas into structured JSON blueprints
Multilingual translation across 46 languages with MiLMMT
Replicate camera motion from a reference video via IC-LoRA
Persian TTS demo using Ava-82M (Kokoro-based)
Relight exterior video clips with a light-direction ball
Define JSON tools and watch Lumma-0.6B generate tool calls
Aerial semantic segmentation with UAVid YOLO26s
Multimodal reasoning VLM with Grug-style thinking
Hyperbolic vision-language zero-shot classification
Multilingual manga speech-bubble OCR (JA/ZH/EN)
4-step RL text-to-image with MeanFlowNFT on SD3.5-Medium
EN subtitle cues to Taiwan Traditional Chinese (0.6B model).
Detect watermarks, signatures, and artifacts in images
Predict robot action chunks from an image + instruction
Generate electron micrographs from text prompts
Forecast Valence & Arousal state change from posts
Graph-native LLM scientific reasoning with graph viz
Speaker-conditioned TTS with emotion & energy control
Video verification & temporal grounding with VideoSearch-R1
MrFlow training-free diffusion acceleration demo
Unified multimodal video generation and editing (5B)
Multi-agent interleaved text-image generation pipeline
P2R fine-grained visual reasoning with Qwen3-VL
VLM-guided tree-search keyframe extraction from videos
Human-object interaction video from image+audio+text
Facial affect estimation with uncertainty via rectified flow
Classify time series in-context with TimEE foundation model
Novel camera viewpoint from a video via depth-warp IC-LoRA
Ultrasound image understanding VLM with MoE architecture
Training-free typographic attack defense for CLIP
Reconstruct video from event-camera voxel grids via LongE2V
4-class lung ultrasound video classifier with attention
Multi-shot cinematic text-to-video generation
Detect AI-generated audio-visual content (DAV-Det)
Spatial reasoning VLM for 3D relations and perspective
Proposal-only typed personal actions with a 35M-param model.
Synthesize contrast-enhanced breast MRI from pre-contrast