Text Generation
PEFT
Safetensors
GGUF
gemma4
unsloth
lora
qlora
fine-tuning
hackathon
gemma-4-good-hackathon
kaggle
translation
speech-recognition
accessibility
on-device
conversational
Instructions to use bradduy/banhmi-gemma4-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bradduy/banhmi-gemma4-e4b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/gemma-4-E4B-it-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "bradduy/banhmi-gemma4-e4b") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bradduy/banhmi-gemma4-e4b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bradduy/banhmi-gemma4-e4b:Q3_K_S # Run inference directly in the terminal: llama cli -hf bradduy/banhmi-gemma4-e4b:Q3_K_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bradduy/banhmi-gemma4-e4b:Q3_K_S # Run inference directly in the terminal: llama cli -hf bradduy/banhmi-gemma4-e4b:Q3_K_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bradduy/banhmi-gemma4-e4b:Q3_K_S # Run inference directly in the terminal: ./llama-cli -hf bradduy/banhmi-gemma4-e4b:Q3_K_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bradduy/banhmi-gemma4-e4b:Q3_K_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf bradduy/banhmi-gemma4-e4b:Q3_K_S
Use Docker
docker model run hf.co/bradduy/banhmi-gemma4-e4b:Q3_K_S
- LM Studio
- Jan
- vLLM
How to use bradduy/banhmi-gemma4-e4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bradduy/banhmi-gemma4-e4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bradduy/banhmi-gemma4-e4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bradduy/banhmi-gemma4-e4b:Q3_K_S
- Ollama
How to use bradduy/banhmi-gemma4-e4b with Ollama:
ollama run hf.co/bradduy/banhmi-gemma4-e4b:Q3_K_S
- Unsloth Desktop
- Docker Model Runner
How to use bradduy/banhmi-gemma4-e4b with Docker Model Runner:
docker model run hf.co/bradduy/banhmi-gemma4-e4b:Q3_K_S
- Lemonade
How to use bradduy/banhmi-gemma4-e4b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bradduy/banhmi-gemma4-e4b:Q3_K_S
Run and chat with the model
lemonade run user.banhmi-gemma4-e4b-Q3_K_S
List all available models
lemonade list
- Atomic Chat
Download apps/macos/scripts/gemma_sidecar.py from bradduy/banhmi-gemma4-e4b: direct link, hf CLI and curl.
- Browser
- Download file 5.71 kB
-
https://huggingface.co/bradduy/banhmi-gemma4-e4b/resolve/main/apps/macos/scripts/gemma_sidecar.py
- Command line
-
hf download hf://bradduy/banhmi-gemma4-e4b/apps/macos/scripts/gemma_sidecar.py
-
curl -L -o gemma_sidecar.py https://huggingface.co/bradduy/banhmi-gemma4-e4b/resolve/main/apps/macos/scripts/gemma_sidecar.py
5.71 kB
| #!/usr/bin/env python3 | |
| """ | |
| Gemma 4 MLX sidecar for the macOS Bánh mì chuyển ngữ app. | |
| Reads NDJSON requests from stdin, one per line. Each request describes an | |
| audio clip (either file path or raw PCM over stdin) and a task. Writes | |
| NDJSON responses to stdout. | |
| Request format (JSON, one per line): | |
| {"task": "transcribe_translate", | |
| "audio_path": "/tmp/clip.wav", | |
| "target_lang": "Vietnamese"} | |
| Response format (JSON, one per line): | |
| {"ok": true, | |
| "source_text": "...", | |
| "translated_text": "...", | |
| "latency_ms": 1234} | |
| On error: | |
| {"ok": false, "error": "..."} | |
| The model is loaded once at startup so subsequent requests are fast. | |
| Usage from Swift: | |
| let proc = Process() | |
| proc.executableURL = URL(fileURLWithPath: "/usr/bin/env") | |
| proc.arguments = ["python3", "/path/to/gemma_sidecar.py"] | |
| // Write one JSON object + newline per request, read back one JSON per line. | |
| """ | |
| import json | |
| import os | |
| import sys | |
| import tempfile | |
| import time | |
| MODEL_ID_DEFAULT = "unsloth/gemma-4-E2B-it-UD-MLX-4bit" | |
| def eprint(*args, **kwargs): | |
| print(*args, file=sys.stderr, flush=True, **kwargs) | |
| def load_model(model_id: str): | |
| from mlx_vlm import load, generate, apply_chat_template | |
| model, processor = load(model_id) | |
| config = getattr(model, "config", None) | |
| return model, processor, config, generate, apply_chat_template | |
| def run_prompt(bundle, audio_path: str, prompt_text: str, max_tokens: int = 256) -> str: | |
| model, processor, config, generate, apply_chat_template = bundle | |
| formatted = apply_chat_template(processor, config, prompt_text, num_audios=1) | |
| result = generate( | |
| model, processor, formatted, | |
| audio=audio_path, | |
| max_tokens=max_tokens, | |
| temperature=0.0, | |
| verbose=False, | |
| ) | |
| return (result.text if hasattr(result, "text") else str(result)).strip() | |
| def handle(req: dict, bundle) -> dict: | |
| task = req.get("task", "transcribe") | |
| audio_path = req.get("audio_path") | |
| target_lang = req.get("target_lang") | |
| max_tokens = int(req.get("max_tokens", 256)) | |
| if not audio_path or not os.path.exists(audio_path): | |
| return {"ok": False, "error": f"audio_path missing or not found: {audio_path!r}"} | |
| # Save each received chunk to /tmp so we can inspect what we actually sent | |
| try: | |
| import shutil | |
| shutil.copy(audio_path, "/tmp/banhmi_last_chunk.wav") | |
| eprint(f"[sidecar] chunk {os.path.getsize(audio_path)} bytes -> /tmp/banhmi_last_chunk.wav") | |
| except Exception as exc: | |
| eprint(f"[sidecar] debug copy failed: {exc}") | |
| t0 = time.time() | |
| if task == "transcribe": | |
| text = run_prompt(bundle, audio_path, "Transcribe this audio", max_tokens) | |
| return { | |
| "ok": True, | |
| "source_text": text, | |
| "translated_text": None, | |
| "latency_ms": int((time.time() - t0) * 1000), | |
| } | |
| if task == "translate": | |
| if not target_lang: | |
| return {"ok": False, "error": "target_lang required for translate task"} | |
| prompt = ( | |
| f"Translate the speech in this audio into {target_lang}. " | |
| f"The speaker may be using any language — detect it and translate. " | |
| f"If the speech is already in {target_lang}, output the speech as-is. " | |
| f"Reply with only the {target_lang} text, no explanations or quotes." | |
| ) | |
| translated = run_prompt(bundle, audio_path, prompt, max_tokens) | |
| return { | |
| "ok": True, | |
| "source_text": None, | |
| "translated_text": translated, | |
| "latency_ms": int((time.time() - t0) * 1000), | |
| } | |
| if task == "transcribe_translate": | |
| if not target_lang: | |
| return {"ok": False, "error": "target_lang required"} | |
| # Two calls: original + translation. Run them sequentially. | |
| translate_prompt = ( | |
| f"Translate the speech in this audio into {target_lang}. " | |
| f"The speaker may be using any language — detect it and translate. " | |
| f"If the speech is already in {target_lang}, output the speech as-is. " | |
| f"Reply with only the {target_lang} text, no explanations or quotes." | |
| ) | |
| translated = run_prompt(bundle, audio_path, translate_prompt, max_tokens) | |
| source = run_prompt(bundle, audio_path, "Transcribe this audio", max_tokens) | |
| return { | |
| "ok": True, | |
| "source_text": source, | |
| "translated_text": translated, | |
| "latency_ms": int((time.time() - t0) * 1000), | |
| } | |
| return {"ok": False, "error": f"unknown task: {task}"} | |
| def main(): | |
| model_id = os.environ.get("GEMMA_MLX_MODEL", MODEL_ID_DEFAULT) | |
| eprint(f"[sidecar] loading {model_id}") | |
| bundle = load_model(model_id) | |
| eprint(f"[sidecar] ready — awaiting requests on stdin") | |
| # Emit a "ready" message so the host knows the model is loaded. | |
| print(json.dumps({"event": "ready", "model": model_id}), flush=True) | |
| # Tempdir for audio uploads if needed | |
| with tempfile.TemporaryDirectory(prefix="gemma_sidecar_") as tmpdir: | |
| for line in sys.stdin: | |
| line = line.strip() | |
| if not line: | |
| continue | |
| try: | |
| req = json.loads(line) | |
| except json.JSONDecodeError as exc: | |
| print(json.dumps({"ok": False, "error": f"invalid JSON: {exc}"}), flush=True) | |
| continue | |
| try: | |
| resp = handle(req, bundle) | |
| except Exception as exc: | |
| resp = {"ok": False, "error": f"{type(exc).__name__}: {exc}"} | |
| print(json.dumps(resp, ensure_ascii=False), flush=True) | |
| if __name__ == "__main__": | |
| main() | |