Instructions to use zveloxy/Veloxy-Flash-15B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zveloxy/Veloxy-Flash-15B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="zveloxy/Veloxy-Flash-15B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("zveloxy/Veloxy-Flash-15B") model = AutoModelForCausalLM.from_pretrained("zveloxy/Veloxy-Flash-15B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use zveloxy/Veloxy-Flash-15B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf zveloxy/Veloxy-Flash-15B:Q4_K_M # Run inference directly in the terminal: llama cli -hf zveloxy/Veloxy-Flash-15B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf zveloxy/Veloxy-Flash-15B:Q4_K_M # Run inference directly in the terminal: llama cli -hf zveloxy/Veloxy-Flash-15B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf zveloxy/Veloxy-Flash-15B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf zveloxy/Veloxy-Flash-15B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf zveloxy/Veloxy-Flash-15B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf zveloxy/Veloxy-Flash-15B:Q4_K_M
Use Docker
docker model run hf.co/zveloxy/Veloxy-Flash-15B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use zveloxy/Veloxy-Flash-15B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zveloxy/Veloxy-Flash-15B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zveloxy/Veloxy-Flash-15B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/zveloxy/Veloxy-Flash-15B:Q4_K_M
- SGLang
How to use zveloxy/Veloxy-Flash-15B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "zveloxy/Veloxy-Flash-15B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zveloxy/Veloxy-Flash-15B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "zveloxy/Veloxy-Flash-15B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zveloxy/Veloxy-Flash-15B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use zveloxy/Veloxy-Flash-15B with Ollama:
ollama run hf.co/zveloxy/Veloxy-Flash-15B:Q4_K_M
- Unsloth Desktop
- Pi
How to use zveloxy/Veloxy-Flash-15B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zveloxy/Veloxy-Flash-15B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "zveloxy/Veloxy-Flash-15B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use zveloxy/Veloxy-Flash-15B with Docker Model Runner:
docker model run hf.co/zveloxy/Veloxy-Flash-15B:Q4_K_M
- Lemonade
How to use zveloxy/Veloxy-Flash-15B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull zveloxy/Veloxy-Flash-15B:Q4_K_M
Run and chat with the model
lemonade run user.Veloxy-Flash-15B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use zveloxy/Veloxy-Flash-15B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zveloxy/Veloxy-Flash-15B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default zveloxy/Veloxy-Flash-15B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use zveloxy/Veloxy-Flash-15B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zveloxy/Veloxy-Flash-15B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "zveloxy/Veloxy-Flash-15B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
⚡ Veloxy-Flash-15B: Faz 8 Tool-Call Master & Autonomous Reasoning Engine ⚡
14.7B Active Parameter Backbone | 38 Native Full-OS & Browser Tools | Dynamic Unknown Tool Discovery | 262K Native Context | Java 21 LTS & Project Loom | Zero-Refusal Cybersecurity | 2026 Temporal Anchor | Full GGUF Multi-Quant Family
📌 1. Overview
Veloxy-Flash-15B (Faz 8: Tool-Call Master) is a next-generation hybrid language model developed by Veloxy AI, engineered specifically for autonomous agent workflows, deep step-by-step reasoning (<think>), 38 native operating system & browser tools, enterprise-grade Java 21 LTS / Project Loom, and uncensored cybersecurity analysis.
Built upon an optimized hybrid architecture alternating between Linear-Attention and Full Multi-Head Attention layers, the model scales smoothly across a native 262,144-token context window with linear memory complexity. Designed for both enterprise cloud deployments and consumer hardware (such as NVIDIA RTX 3060 12GB, RTX 4060/4070, and RTX 5090), it delivers zero-latency, high-precision local execution.
🌟 2. Key Capabilities in Faz 8
Faz 8 Master training was conducted across 4,342 curated multi-turn enterprise samples on an NVIDIA RTX A6000 48GB cluster, achieving a record convergence loss of 0.2799.
🛠️ 1. 38 Native Full-OS & Browser Tools (<tool_call>)
The model natively produces strict JSON parameters for 38 system-level tools covering the terminal, filesystem, Git version control, DOM automation, and security diagnostics without requiring external wrappers.
🔍 2. Dynamic Unknown Tool Discovery
When faced with an unfamiliar command or custom user tool (veloxy_kernel_probe, custom_mcp_bridge, etc.), the model avoids hallucinations. It autonomously applies introspection strategies: passing --help, querying available schemas, and inferring valid arguments on the fly.
🔓 3. Zero-Refusal Cybersecurity & Binary Exploitation
Engineered for white-hat security researchers, penetration testers, and reverse engineers with technical depth and zero artificial refusals:
- Binary Exploitation: ASLR/DEP evasion techniques, GOT/PLT overwriting, ROP chain synthesis, stack/heap canary analysis.
- Web & Network Security: SAST source code auditing, CVE analysis, SSRF, insecure deserialization, SQLi, and IDOR vulnerability research with actionable remediation.
☕ 4. Elite Java 21 LTS & JVM Internals
- Project Loom: Virtual Threads (
Executors.newVirtualThreadPerTaskExecutor()), Structured Concurrency, and Scoped Values for high-throughput concurrency. - JVM Internals: ZGC (Generational), Shenandoah, and G1 GC mechanics, JIT C1/C2 compilation stages, Java Memory Model (JMM) barriers, and bytecode disassembly.
- Spring Boot 3: Reactive WebFlux pipelines, GraalVM AOT Native Image optimizations.
⏳ 5. 2026 Temporal Anchor
The model’s temporal baseline is strictly fixed to 2026. It assumes modern 2026 ecosystem baselines (Java 21/22, Next.js 15, React 19, Python 3.12+, PyTorch 2.6) by default and does not hallucinate outdated legacy patterns.
🧠 6. Scalable Reasoning Depth (Reasoning Effort)
Reasoning depth can be dynamically controlled via prompts across 5 tiers:
[EFFORT: LOW]– Fast, direct responses.[EFFORT: MEDIUM]– Balanced daily reasoning.[EFFORT: HIGH]– In-depth verification and step-by-step math.[EFFORT: XHIGH]– Olympiad math, complex algorithm design, and proofs.[EFFORT: MAX]– Full architectural breakdown, root-cause analysis, and exhaustive code audits.
🧰 3. 38 Native Tool Catalog
| Category | Native Tools | Description |
|---|---|---|
| System & Terminal | bash_exec, terminal_spawn, process_list, process_kill, service_status, service_restart, cron_schedule |
Complete OS process and daemon management |
| Filesystem | file_read, file_write, file_diff_patch, file_search, dir_tree, dir_create, file_delete |
Atomic reads, writes, patches, and directory trees |
| Developer & Git | git_status, git_diff, git_commit, git_branch, code_ast_parse, linter_run, test_runner |
Autonomous version control, AST parsing, and testing |
| Browser & Web | browser_open, element_click, input_type, page_extract, screenshot_capture, network_har_inspect |
Full-DOM interaction, screenshotting, and HAR analysis |
| Memory & Agents | memory_store, memory_retrieve, session_summary, task_decompose, subagent_invoke |
Long-term memory storage and subagent delegation |
| Security & Diagnostics | cve_search, binary_disasm, network_port_scan, sast_code_audit, heap_inspector |
Vulnerability assessment, disassembly, and port scanning |
📦 4. Multi-Quant GGUF Family
All primary GGUF quantizations are available directly in this repository for local use:
| Quantization | File Size | Required VRAM | Recommended Hardware | Quality / PPL | Recommended Use Case |
|---|---|---|---|---|---|
veloxy-flash-15b-Q8_0.gguf |
9.53 GB | ~12-16 GB | RTX 4080, RTX 3090, RTX 4090 | 100% (Pristine 8-bit) | Maximum precision benchmark reference |
veloxy-flash-15b-Q6_K.gguf |
6.85 GB | ~9-12 GB | RTX 3060 12GB, RTX 4070 | 99.8% | High-fidelity enterprise production |
veloxy-flash-15b-Q5_K_M.gguf |
6.02 GB | ~8-10 GB | RTX 3060 12GB, RTX 4060 Ti 16GB | 99.2% | Balanced speed and memory efficiency |
veloxy-flash-15b-Q4_K_M.gguf ⭐ |
5.24 GB | ~6-8 GB | NVIDIA RTX 3060 12GB | 98.5% (Recommended) | 100% GPU offload on 12GB VRAM cards |
veloxy-flash-15b-Q3_K_M.gguf |
4.40 GB | ~6 GB | RTX 2060 6GB, RTX 3050 | 96.0% | Lower VRAM environments |
veloxy-flash-15b-Q2_K.gguf |
3.83 GB | ~4-6 GB | GTX 1660, Apple Silicon 8GB | 91.5% | Ultra-light edge computing devices |
💡 Notice for RTX 3060 12GB Users: The
Q4_K_Mquant consumes only ~5.6 GB VRAM in memory, leaving ~6.5 GB free headroom for system and KV-Cache.
(Community i-Matrix / IQ quants are also provided by mradermacher/Veloxy-Flash-15B-i1-GGUF.)
🚀 5. Quickstart & Deployment Guide
A. Run with Ollama (Recommended for Local PC)
- Download
veloxy-flash-15b-Q4_K_M.ggufinto your directory. - Create a
Modelfile:
FROM ./veloxy-flash-15b-Q4_K_M.gguf
PARAMETER num_ctx 262144
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
PARAMETER temperature 0.6
PARAMETER top_p 0.9
SYSTEM """You are Veloxy-Flash-15B, an advanced 15-billion parameter artificial intelligence model developed by Veloxy AI.
- Capabilities: Deep reasoning (<think>), 38 native OS/Browser tools, Java 21+ and JVM internals, and zero-refusal cybersecurity research.
- Temporal Anchor: Current year is 2026.
- Reasoning Effort: Scale your thinking inside <think> tags according to task complexity."""
- Build and launch:
ollama create veloxy-flash-15b -f Modelfile
ollama run veloxy-flash-15b
B. Python / Transformers (Native bfloat16 & 4-bit)
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
model_id = "zveloxy/Veloxy-Flash-15B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
# 4-bit quantization for 12GB GPUs (RTX 3060 / 4060):
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=bnb_config,
device_map="auto",
trust_remote_code=True
)
messages = [
{"role": "system", "content": "You are Veloxy-Flash-15B, an autonomous reasoning and systems engineering agent developed by Veloxy AI."},
{"role": "user", "content": "Design an eBPF program in C that hooks onto sys_enter_execve to detect suspicious reverse shell invocations."}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([prompt], return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.6)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
C. llama.cpp Server (OpenAI-Compatible Local API)
llama-server \
-m veloxy-flash-15b-Q4_K_M.gguf \
-c 32768 \
-ngl 99 \
--port 8080 \
--host 0.0.0.0
🛠️ 6. Tool-Calling Format Example
<|im_start|>user
Inspect port 8443 on the staging server and read the nginx reverse proxy configuration.
<|im_end|>
<|im_start|>assistant
<think>
1. Check port 8443 status using network_port_scan.
2. Read the nginx config file at /etc/nginx/conf.d/proxy.conf.
</think>
<tool_call>
<function=network_port_scan>
<parameter=host>192.168.1.50</parameter>
<parameter=ports>[8443]</parameter>
</function>
</tool_call>
<tool_call>
<function=file_read>
<parameter=path>/etc/nginx/conf.d/proxy.conf</parameter>
</function>
</tool_call><|im_end|>
🌐 7. Veloxy AI Ecosystem
| Model | Format / Parameters | Focus & Target Hardware | Link |
|---|---|---|---|
| Veloxy-Flash-15B ⭐ | 15B BF16 + GGUF Family | Faz 8 Tool-Call Master, 38 Native Tools, RTX 3060 12GB & Edge | HF Repo |
| Veloxy-Flash-15B-LoRA 🧩 | PEFT Rank 16 Adapter | Official Faz 8 Integration Adapter Weights | HF Repo |
| Veloxy-Titan-32B 🛸 | 32B BF16 / 1M Context | Flagship Deep Reasoning, 1M Context, H100 NVL Benchmark | HF Repo |
| Veloxy-Titan-32B-Vision 👁️ | 32B + 27-Layer ViT | Multimodal Vision & Reasoning, UI-to-Code, Schematics | HF Repo |
📄 8. License
Veloxy-Flash-15B is released under the Apache License 2.0. Free for academic, commercial, and private deployment.
- Downloads last month
- 721
Model tree for zveloxy/Veloxy-Flash-15B
Base model
Qwen/Qwen3.5-9B-BaseEvaluation results
- Final Train Loss on Faz 8 Master Enterprise Benchmarkself-reported0.280
- Tool-Calling Schema Precision on Faz 8 Master Enterprise Benchmarkself-reported99.4%
- Context Window Length on Faz 8 Master Enterprise Benchmarkself-reported262144.000