Instructions to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with llama-cpp-python:

# !pip install llama-cpp-python

from llama_cpp import Llama

llm = Llama.from_pretrained(
	repo_id="magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF",
	filename="granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0.gguf",
)

llm.create_chat_completion(
	messages = [
		{
			"role": "user",
			"content": "What is the capital of France?"
		}
	]
)

Notebooks
Google Colab
Kaggle
Local Apps

llama.cpp

How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with llama.cpp:

Install from brew

brew install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama-server -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0
# Run inference directly in the terminal:
llama-cli -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Install from WinGet (Windows)

winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama-server -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0
# Run inference directly in the terminal:
llama-cli -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Use pre-built binary

# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0
# Run inference directly in the terminal:
./llama-cli -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Build from source code

git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0
# Run inference directly in the terminal:
./build/bin/llama-cli -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Use Docker

docker model run hf.co/magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

LM Studio
Jan

vLLM

How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Ollama
How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with Ollama:
```
ollama run hf.co/magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0
```

Unsloth Studio new

How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with Unsloth Studio:

Install Unsloth Studio (macOS, Linux, WSL)

curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF to start chatting

Install Unsloth Studio (Windows)

irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF to start chatting

Using HuggingFace Spaces for Unsloth

# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF to start chatting

Pi new

How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with Pi:

Start the llama.cpp server

# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama-server -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Configure the model in Pi

# Install Pi:
npm install -g @mariozechner/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "llama-cpp": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0"
        }
      ]
    }
  }
}

Run Pi

# Start Pi in your project directory:
pi

Hermes Agent new

How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with Hermes Agent:

Start the llama.cpp server

# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama-server -hf magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Configure Hermes

# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Run Hermes

hermes

Docker Model Runner
How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with Docker Model Runner:
```
docker model run hf.co/magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0
```

Lemonade

How to use magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF with Lemonade:

Pull the model

# Download Lemonade from https://lemonade-server.ai/
lemonade pull magiccodingman/Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF:Q8_0

Run and chat with the model

lemonade run user.Granite-4.0-H-350M-Unsloth-MagicQuant-Hybrid-GGUF-Q8_0

List all available models

lemonade list

magiccodingman commited on Dec 2, 2025

Commit

7f47feb

verified ·

1 Parent(s): 6e75d88

File name changes

Browse files

Files changed (4) hide show

.gitattributes +2 -2
README.md +18 -21
granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0.gguf +3 -0
granite-4.0-h-350m-unsloth-mxfp4_moe-O-Q6K-EQKUD-Q8_0.gguf +3 -0

.gitattributes CHANGED Viewed

@@ -33,5 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
-granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD=B16-O=Q6K-Q=Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
-granite-4.0-h-350m-unsloth-mxfp4_moe-O=Q6K-EQKUD=Q8_0.gguf filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
+granite-4.0-h-350m-unsloth-mxfp4_moe-O-Q6K-EQKUD-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text

README.md CHANGED Viewed

@@ -35,24 +35,22 @@ To dive deeper into how MagicQuant works, see the main repo:
 | model_name | file_size_gb | bench_tps | avg_prec_loss |
 | ---------- | ------------ | --------- | ------------- |
-| [mxfp4_moe-EKUD=B16-O=Q6K-Q=Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD=B16-O=Q6K-Q=Q8_0.gguf?download=true) | 0.54 | 1705.35 | 0.0816 |
-| [mxfp4_moe-O=Q6K-EQKUD=Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-O=Q6K-EQKUD=Q8_0.gguf?download=true) | 0.34 | 1605.97 | 0.2555 |
 ### Table - PPL Columns
 | model_name | gen | gen_er | code | code_er | math | math_er |
 | ---------- | --- | ------ | ---- | ------- | ---- | ------- |
-| [mxfp4_moe-EKUD=B16-O=Q6K-Q=Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD=B16-O=Q6K-Q=Q8_0.gguf?download=true) | 18.1560 | 0.4667 | 1.9548 | 0.0175 | 10.2986 | 0.2319 |
-| [mxfp4_moe-O=Q6K-EQKUD=Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-O=Q6K-EQKUD=Q8_0.gguf?download=true) | 18.2304 | 0.4691 | 1.9555 | 0.0175 | 10.3074 | 0.2320
 ### Table - Precision Loss Columns
 | model_name | loss_general | loss_code | loss_math |
 | ---------- | ------------ | --------- | --------- |
-| [mxfp4_moe-EKUD=B16-O=Q6K-Q=Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD=B16-O=Q6K-Q=Q8_0.gguf?download=true) | 0.1368 | 0.0051 | 0.1030 |
-| [mxfp4_moe-O=Q6K-EQKUD=Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-O=Q6K-EQKUD=Q8_0.gguf?download=true) | 0.5471 | 0.0307 | 0.1886 |
 ---
@@ -88,13 +86,13 @@ This tells users the *default* quantization for the majority of tensors.
 If certain tensor groups use a different quant scheme than the base, they appear afterwards as:
 ```
-<GroupLetters>=<Quant>
 ```
 Multiple group blocks can be chained:
 ```
-<Model>-<Base>-<Groups>=<Quant>-<Groups>=<Quant>...
 ```
 Only *exceptions* appear.
@@ -128,7 +126,7 @@ These are the compact codes for each major tensor group:
 If multiple groups share the same quant scheme, combine them:
 ```
-EH=B16
 ```
 Means:
@@ -139,7 +137,7 @@ Means:
 Another block:
 ```
-QKO=IQ4NL
 ```
 Means:
@@ -163,7 +161,7 @@ You can stack as many blocks as needed.
 ### **Hybrid Name:**
 ```
-Qwen3-4B-MXFP4-EH=B16-QKO=IQ4NL.gguf
 ```
 This reads as:
@@ -182,13 +180,13 @@ Clean. Simple. Understandable. Portable.
 Example: Only embeddings become Q5_K.
 ```
-Qwen3-4B-MXFP4-E=Q5K.gguf
 ```
 If only MoE router changes:
 ```
-Qwen3-4B-MXFP4-R=Q8.gguf
 ```
 ---
@@ -207,8 +205,7 @@ No extra suffixes.
 ## 🧠 **8. Notation Guidelines (For Clarity and Aesthetics)**
-* Use hyphens `-` between blocks.
-* Use equals `=` between group and quant scheme.
 * No need for `_` unless the quant type requires it (`IQ4_NL`, `Q4_K_M`, etc.).
 * Order of groups doesn’t matter, but a consistent order is recommended:
@@ -222,11 +219,11 @@ This mirrors information flow: embeddings → attention → ffn → moe.
 | Description                                | Final Name                     |
 | ------------------------------------------ | ------------------------------ |
-| Base MXFP4, only embeddings = BF16         | `MXFP4-E=B16`                  |
-| Base IQ4_NL, Q/K/O = Q6_K                  | `IQ4NL-QKO=Q6K`                |
-| Base Q6_K, Head & Router = BF16            | `Q6K-HR=B16`                   |
 | Everything BF16 → no hybrid                | `B16` (or `F16/BF16` baseline) |
-| Full MoE override: experts + router = Q8_0 | `X R=Q8_0` → `XR=Q8_0`         |
 ---

 | model_name | file_size_gb | bench_tps | avg_prec_loss |
 | ---------- | ------------ | --------- | ------------- |
+| [mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0.gguf?download=true) | 0.54 | 1705.35 | 0.0816 |
+| [mxfp4_moe-O-Q6K-EQKUD-Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-O-Q6K-EQKUD-Q8_0.gguf?download=true) | 0.34 | 1605.97 | 0.2555 |
 ### Table - PPL Columns
 | model_name | gen | gen_er | code | code_er | math | math_er |
 | ---------- | --- | ------ | ---- | ------- | ---- | ------- |
+| [mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0.gguf?download=true) | 18.1560 | 0.4667 | 1.9548 | 0.0175 | 10.2986 | 0.2319 |
+| [mxfp4_moe-O-Q6K-EQKUD-Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-O-Q6K-EQKUD-Q8_0.gguf?download=true) | 18.2304 | 0.4691 | 1.9555 | 0.0175 | 10.3074 | 0.2320
 ### Table - Precision Loss Columns
 | model_name | loss_general | loss_code | loss_math |
 | ---------- | ------------ | --------- | --------- |
+| [mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0.gguf?download=true) | 0.1368 | 0.0051 | 0.1030 |
+| [mxfp4_moe-O-Q6K-EQKUD-Q8_0](./../../resolve/main/granite-4.0-h-350m-unsloth-mxfp4_moe-O-Q6K-EQKUD-Q8_0.gguf?download=true) | 0.5471 | 0.0307 | 0.1886 |
 ---
 If certain tensor groups use a different quant scheme than the base, they appear afterwards as:
 ```
+<GroupLetters>-<Quant>
 ```
 Multiple group blocks can be chained:
 ```
+<Model>-<Base>-<Groups>-<Quant>-<Groups>-<Quant>...
 ```
 Only *exceptions* appear.
 If multiple groups share the same quant scheme, combine them:
 ```
+EH-B16
 ```
 Means:
 Another block:
 ```
+QKO-IQ4NL
 ```
 Means:
 ### **Hybrid Name:**
 ```
+Qwen3-4B-MXFP4-EH-B16-QKO-IQ4NL.gguf
 ```
 This reads as:
 Example: Only embeddings become Q5_K.
 ```
+Qwen3-4B-MXFP4-E-Q5K.gguf
 ```
 If only MoE router changes:
 ```
+Qwen3-4B-MXFP4-R-Q8.gguf
 ```
 ---
 ## 🧠 **8. Notation Guidelines (For Clarity and Aesthetics)**
+* Use hyphens `-` throughout the naming scheme for consistency.
 * No need for `_` unless the quant type requires it (`IQ4_NL`, `Q4_K_M`, etc.).
 * Order of groups doesn’t matter, but a consistent order is recommended:
 | Description                                | Final Name                     |
 | ------------------------------------------ | ------------------------------ |
+| Base MXFP4, only embeddings = BF16         | `MXFP4-E-B16`                  |
+| Base IQ4_NL, Q/K/O = Q6_K                  | `IQ4NL-QKO-Q6K`                |
+| Base Q6_K, Head & Router = BF16            | `Q6K-HR-B16`                   |
 | Everything BF16 → no hybrid                | `B16` (or `F16/BF16` baseline) |
+| Full MoE override: experts + router = Q8_0 | `X R-Q8_0` → `XR-Q8_0`         |
 ---

granite-4.0-h-350m-unsloth-mxfp4_moe-EKUD-B16-O-Q6K-Q-Q8_0.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:689e694797a2074350577fea7ad61a8167c6e0ee98bfbbfb840c28ac76974310
+size 580910016

granite-4.0-h-350m-unsloth-mxfp4_moe-O-Q6K-EQKUD-Q8_0.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:5e5a53b4e80abe4039661da23e477a3b8f4b641b7b40a70829c0243db0dd437f
+size 365624256