gemma4
Here are 352 public repositories matching this topic...
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
-
Updated
Jul 31, 2026 - Swift
A powerful Zotero AI and MCP plugin with ChatGPT, Gemini 3.6, Claude Fable 5, Claude Sonnet 5, DeepSeek V4, Grok, OpenRouter, Kimi k3, GLM 5.2, SiliconFlow, GPT-oss, Gemma 4, Qwen 3.7
-
Updated
Jul 30, 2026 - JavaScript
Steno is the privacy-first AI notepad & notetaker for all your confidential conversations. On Windows & MacOS. Perfect for government, healthcare, defence, legal and CXOs.
-
Updated
Jul 31, 2026 - TypeScript
OpenClaw alternative in your pocket
-
Updated
Jul 31, 2026 - Kotlin
PokeClaw (PocketClaw) — first on-device AI that controls your Android phone. Gemma 4, no cloud, no API key. Poke is short for Pocket.
-
Updated
Jun 2, 2026 - Kotlin
Plug-and-play local AI studio: uncensored chat, image & video generation, coding agent. Runs abliterated LLMs + ComfyUI 100% offline. One installer, no Docker, no cloud.
-
Updated
Jul 31, 2026 - TypeScript
Gemma Gem runs Google's Gemma 4 model entirely on-device via WebGPU — no API keys, no cloud, no data leaving your machine.
-
Updated
May 29, 2026 - TypeScript
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
-
Updated
Aug 1, 2026 - Python
Open-source AI browser agent for Chrome and Firefox (monorepo) 🧠
-
Updated
Jul 31, 2026 - JavaScript
An open-source Cotypist with macOS system wide AI autocomplete
-
Updated
Jun 27, 2026 - Swift
Community model zoo for Apple Core AI (iOS/macOS 27): 57 models — LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting — each gated against its source model and shipped with the recipe that produced it. Downloadable from Hugging Face, runnable in one line of Swift via CoreAIKit. Plus benchmarks, Metal kernels, knowledge base.
-
Updated
Jul 31, 2026 - Python
Run local LLMs like Gemma, Qwen, and LLaMA on Android for offline, private, real-time chat and question answering with LiteRT and ONNX Runtime.
-
Updated
May 6, 2026 - Kotlin
FFPA: Kernel Library for Large Headdim Attention - 1.5x~6x speedup over PyTorch SDPA.
-
Updated
Jul 31, 2026 - Python
llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).
-
Updated
Jul 30, 2026 - C++
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability
-
Updated
Jul 31, 2026 - C#
This is end to end course on AI Agents and Agentic AI with 15+ AI Agent Projects with real time use cases and industry expertise.
-
Updated
Apr 17, 2026 - Jupyter Notebook
Private on-device AI chat for Android — runs any GGUF model locally via llama.cpp with ARM-optimised SIMD. Zero network permissions, encrypted settings, biometric lock, tamper detection. + GPU Acceleration
-
Updated
Jul 10, 2026 - Kotlin
Agentic ✧ Gemma Inference for Android System Intelligence
-
Updated
Jul 25, 2026 - Kotlin
Improve this page
Add a description, image, and links to the gemma4 topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the gemma4 topic, visit your repo's landing page and select "manage topics."