✨ We are happy to share with you our new universal LLM models based on Qwen3 1.7B and 4B — powerful, multilingual and ready to solve a wide range of problems!
🛠️ We have conducted additional training and carefully merged them to achieve even better results and maximize the potential of the models.
🆓 And most importantly — the models are completely open and free under the Apache-2.0 license!
Hey all Finally it's happening. DeepGit lite is back now, running on cpu only devices. Just smartly search across Github and spin up conversational agents in the background and have grounded conversation with repositories Try it out now!!!! zamal/DeepGit
Say hallo to GermaNER 💪– a lightweight, high-accuracy NER model for German texts, powered by XLM-RoBERTa + LoRA adapters! ⚡ Fast, efficient, and open-source – perfect for tagging names, places & orgs in real-world German data. Try it now on Hugging Face 👉 fau/GermaNER
🚀 Videoxity is live on Hugging Face! 🎞️ A powerful, modular toolkit for intelligent video manipulation and scene editing.
With Videoxity, you can:
🖼️ Auto-caption keyframes with BLIP
🧠 Filter scenes using natural language (e.g. “remove dog scenes”)
✂️ Seamlessly trim videos with FFmpeg
📊 Generate frame-based summaries
Powered by Groq LLM + LangChain, OpenCV, BLIP, and SentenceTransformers, Videoxity bridges vision and language to give developers full control over video content. 🔧 Built for developers. Feedback welcome!
Hey folks! Just launched DeepGit Lite — a lighter version of DeepGit with fewer components under the hood. It won’t perform quite like the full powerhouse, but it’s great for a quick peek and first-hand feel! ⚙️👀
DeepGit: Your GitHub Gold Digger! 💰🚀 Hey Hugging Face gang! Meet DeepGit—my open-source sidekick that rips through GitHub to snag repos that fit you. Done with dead-end searches? Me too. Built it with LangGraph and some dope tricks: Embeddings grab the good stuff (HF magic, baby!)
Re-ranking nails the best picks
Snoops docs, code, and buzz in one slick flow
Drops a clean list of hidden gems 💎
Unearth that sneaky ML lib or Python gem—run python app.py or langgraph dev and boom! Peek it at https://github.com/zamalali/DeepGit. Fork it, tweak it, love it—Docker’s in, HF vibes are strong. Drop a 🌟 or a crazy idea—I’m pumped to jam with you all! 🪂
🚀 ftBoost is LIVE – Stop Struggling with Fine-Tuning Data!
Alright folks, if you’re tired of manually crafting fine-tuning datasets, ftBoost is here to do the heavy lifting. One-click, LangChain-Groq-powered data augmentation that scales your training data in OpenAI, Gemini, Mistral, and LLaMA formats—automatically.
🔥 What’s inside? ✅ Smart Augmentations – Paraphrasing, back translation, synonym swapping & synthetic noise. ✅ No more JSONL headaches – Auto-formats everything for OpenAI, Gemini, Mistral & LLaMA. ✅ Custom tuning – Adjust similarity, diversity, and fluency in real-time. ✅ Upload, generate, download – That’s it.
⚡ If you’re fine-tuning LLMs, this will save you hours.
Introducing our first standalone model – FluentlyLM Prinum
Introducing the first standalone model from Project Fluently LM! We worked on it for several months, used different approaches and eventually found the optimal one.
General characteristics: - Model type: Causal language models (QwenForCausalLM, LM Transformer) - Number of parameters: 32.5B - Number of parameters (not embedded): 31.0B - Number of layers: 64 - Context: 131,072 tokens - Language(s) (NLP): English, French, Spanish, Russian, Chinese, Japanese, Persian (officially supported) - License: MIT
Creation strategy: The basis of the strategy is shown in Pic. 2. We used Axolotl & Unsloth for SFT-finetuning with PEFT LoRA (rank=64, alpha=64) and Mergekit for SLERP and TIES mergers.
Interact with your PDF documents like never before! 🤯 Extract text & images, then ask context-aware questions based on both. Powered by RAG techniques & multimodal LLMs. Perfect for studying, research & more! 📝👀 Try it out now!!!! ✍️
✒️ Ultraset - all-in-one dataset for SFT training in Alpaca format. fluently-sets/ultraset
❓ Ultraset is a comprehensive dataset for training Large Language Models (LLMs) using the SFT (instruction-based Fine-Tuning) method. This dataset consists of over 785 thousand entries in eight languages, including English, Russian, French, Italian, Spanish, German, Chinese, and Korean.
🤯 Ultraset solves the problem faced by users when selecting an appropriate dataset for LLM training. It combines various types of data required to enhance the model's skills in areas such as text writing and editing, mathematics, coding, biology, medicine, finance, and multilingualism.
🤗 For effective use of the dataset, it is recommended to utilize only the "instruction," "input," and "output" columns and train the model for 1-3 epochs. The dataset does not include DPO or Instruct data, making it suitable for training various types of LLM models.
❇️ Ultraset is an excellent tool to improve your language model's skills in diverse knowledge areas.