Run local LLMs like Gemma, Qwen, and LLaMA on Android for offline, private, real-time chat and question answering with LiteRT and ONNX Runtime.
-
Updated
May 6, 2026 - Kotlin
Run local LLMs like Gemma, Qwen, and LLaMA on Android for offline, private, real-time chat and question answering with LiteRT and ONNX Runtime.
Training pipeline for an osu!mania 7k next-event model, designed as the upstream predictor for audio-driven 7k map generation systems.
run ai without internet on web and mobile
Multimodal BBQ (bias) visual question answering: pick the right answer from image + context + question + 3 options, and answer "unknown" when evidence is insufficient. Offline, open-weights Qwen2.5-VL (4-bit) + balanced chain-of-thought. Public Balanced Accuracy 0.883.
A Cloud-to-Edge MLOps pipeline for offline industrial diagnostics. Fine-tunes Phi-3-mini (3.8B) on Cloud GPUs via QLoRA, quantizes to INT4, and deploys as a CPU-optimized ONNX microservice for industrial standard sensor logs.
Offline Android flower identification using Kotlin, Jetpack Compose, and a 102-class TensorFlow Lite MobileNetV2 model.
Event-driven image processing pipeline with automatic format normalization, thumbnail generation, AI-powered captioning (BLIP), and full lifecycle management. Built with .NET 10, Python, Docker, and Terraform.
Download Hugging Face model snapshots locally for offline inference, deployment, and Docker workflows.
A comprehensive toolkit for streamlining and simplifying the offline inference process for LLMs across various models and libraries.
Offline CrowdAware system for Raspberry Pi 4B and Heltec LoRa V3 using Raspberry Pi Camera Module 3 and MLX90640 Thermal Camera.
Offline Krita plugin for e621 image tag generation using JTP-3 ONNX Runtime CPU inference
GPT-OSS B20 Local Execution. Lightweight local environment for running it with Python 3.12 and CUDA acceleration. - Run GPT-OSS B20 entirely offline - Optimize text generation with GPU - Enable fast, secure inference on consumer hardware.
Мультимодальная офлайновая система детекции контрафакта (текст+изображение+таблица).
Inclusive hand gesture recognition system for assistive human–computer interaction, based on classical machine learning and MediaPipe Hands.
Aptus-R: An offline candidate ranking system using dual retrieval (FAISS + BM25), a 5-signal composite, and local Phi-3-mini reranking.
Real-time semantic audio codec achieving 300bps bandwidth via gen AI reconstruction.
Add a description, image, and links to the offline-inference topic page so that developers can more easily learn about it.
To associate your repository with the offline-inference topic, visit your repo's landing page and select "manage topics."