Skip to content

Repository files navigation

moe-on-arc

moe-on-arc: Mixture-of-Experts on integrated graphics By Adam Milton-Barker - CogniTech Systems Ltd. Intel Software Innovator.

moe-on-arc is a benchmark harness for measuring accuracy and throughput together when serving a local LLM, plus a verified end-to-end configuration for Qwen3.6-35B-A3B (35B parameters, 3B active, INT4) on an Intel Arc 140V iGPU via OpenVINO Model Server.

The Inspiration

Intel published figures for this model on an AI PC: 42-43 tok/s on a Core Ultra 7 368H, measured with the LLM Bench tool from openvino.genai, in-process via VLMPipeline on OpenVINO 2026.2 (article).

That post is what convinced me this was worth trying. I had a 258V rather than a 368H, and I wanted the model as a service rather than a library, since anything I would build against it needs an HTTP endpoint. Serving through OVMS also adds continuous batching, prefix caching and scheduling that in-process inference does not have.

So the numbers here are not comparable to Intel's: different silicon, different serving path, different runtime version. Intel's post notes the same thing, that a benchmark is a runtime version, a model artifact, a device, a prompt shape and an attention path. This repo is one more configuration on that list, on hardware not covered by the published figures, with accuracy measured alongside throughput.

Install

git clone https://github.com/CogniTechSystems/moe-on-arc
cd moe-on-arc
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python check_env.py

check_env.py verifies versions, enumerates devices with their real memory properties, and checks headroom. Run it before downloading ~18 GB.

Qwen3.6 requires Transformers 5.2, which is why requirements.txt pins it. It is not pulled in by the OpenVINO or Optimum installs, so it must be installed explicitly.

Setup

1. Raise the shared GPU memory limit

Shared GPU Memory Override in Intel Graphics Software

Windows exposes about half of system RAM as the graphics budget, 16 GB on a 32 GB machine. The weights are 18.3 GB, so the default configuration fails with CL_OUT_OF_RESOURCES.

Intel Graphics Software → Graphics → General → Shared GPU Memory Override. Raise it, then confirm:

$ python check_env.py
--- versions ---
[ ok ] openvino 2026.3.0-22451-8a17657b995-releases/2026/3
[ ok ] openvino-genai 2026.3.0.0-3277-bd8d6542e3c
[ ok ] transformers 5.2.0
--- devices ---
       enumerated: CPU, GPU, NPU
[ ok ] GPU: Intel(R) Arc(TM) 140V GPU (16GB) (iGPU)
       (the size in that name is a static driver string, not the live budget)
       total addressable          22.1 GB
       raise via Intel Graphics Software > Graphics > General > Shared GPU Memory Override
[ ok ] NPU present (GPU is the recommended device for Qwen3.6-35B-A3B)
--- memory ---
[ ok ] system RAM 31.6 GB total
[warn] only 21.9 GB free - close things before loading
       INT4 weights are ~19 GB and the iGPU shares this pool
--- disk ---
[ ok ] 484.9 GB free on this volume (model download needs ~20 GB)

Environment looks good. Next: download the model (see README.md).

22.1 GB against 18.3 GB of weights leaves room for the KV cache.

2. Download the model

hf download OpenVINO/Qwen3.6-35B-A3B-int4-ov --local-dir ov_models/qwen3.6-35b-a3b-int4-ov

This artifact is already converted to OpenVINO IR and quantized to INT4. No conversion step is needed.

3. Serve with OVMS

Download ovms_windows_<version>_python_off.zip from the model_server releases. The embedded interpreter is only needed for custom pipeline nodes.

$ovms = "C:\ovms\ovms\ovms.exe"

& $ovms --rest_port 8000 `
  --model_path ov_models\qwen3.6-35b-a3b-int4-ov `
  --model_name qwen3.6-35b-a3b `
  --task text_generation `
  --target_device GPU `
  --tool_parser hermes3 `
  --cache_dir .ov_cache `
  --enable_prefix_caching true
Flag Why
--tool_parser hermes3 Parses tool calls out of model output. Required for agentic use.
--cache_dir Caches the compiled graph. First load ~64 s, later loads much faster.
--enable_prefix_caching Avoids re-prefilling a repeated system prompt. Valuable given weak prefill.

This is the single-model form, so no config.json is needed; that file is only for serving several models, and --add_to_config requires a valid one to already exist.

Do not pass --max_prompt_len on GPU. It is NPU-only and fails with Option not found: MAX_PROMPT_LEN.

Wait for state changed to: AVAILABLE. The REST port binds before loading finishes, so an empty /v3/models early on is expected.

4. Verify

The OVMS endpoint path is /v3, not /v1. We will write the body to a file as PowerShell mangles inline JSON escaping, and -Encoding ascii avoids a BOM the parser rejects.

'{"model":"qwen3.6-35b-a3b","messages":[{"role":"user","content":"Reply with exactly: working"}],"max_tokens":20,"chat_template_kwargs":{"enable_thinking":false}}' |
  Set-Content req.json -Encoding ascii

curl.exe http://localhost:8000/v3/chat/completions -H "Content-Type: application/json" -d "@req.json"

5. Benchmark

Leave OVMS running in its window. In a second terminal, with the venv active:

copy targets.example.json targets.json
python moe_on_arc.py -t task_invoices.json -T targets.json --only q36-ovms -r 3 --max-tokens 512 --failures --save results.json

Usage

python moe_on_arc.py -t TASK -T TARGETS [options]
Flag Effect
-t, --task Task JSON file (required)
-T, --targets Targets JSON file (required)
-r, --repeat N Runs per case; enables variance reporting
--only LABEL Run a subset (repeatable)
--no-schema Disable constrained decoding
--failures Print failing cases
--save FILE Write full results as JSON
--max-tokens N Generation cap (default 256)
--timeout SECONDS Per-request timeout (default 300)
--quiet Suppress the progress counter

A target is abandoned after two consecutive errors. A GPU OOM can leave the driver context in a state where further calls hang.

Task format

{
  "name": "invoice extraction",
  "system": "You extract structured data from invoice text...",
  "template": "Extract the fields from this invoice:\n\n{text}",
  "numeric_tolerance": 0.001,
  "schema": { "type": "object", "properties": {}, "required": [] },
  "cases": [ { "text": "raw input", "expected": { "field": "value" } } ]
}

Every key in expected is graded independently. Numbers use numeric_tolerance (relative) and tolerate currency formatting; strings are compared case- and whitespace-insensitively; lists order-insensitively.

task_invoices.json holds ten invoice cases built to break things: a credit note with a negative total, a PO number adjacent to the invoice number, subtotals directly above grand totals, and dates in ISO, DD/MM/YYYY, Norwegian dotted and 31st January forms.

The harness is domain-agnostic. Swap in log lines, CVs, tickets, listings, anything with a known answer.

Target format

targets.json is a list of objects, each with a label and a backend.

Backend Transport Works with
openai /v1/chat/completions OVMS (on /v3), LM Studio, vLLM, llama-server, mlx_lm.server
ovgenai in-process OpenVINO GenAI on Intel CPU, iGPU, NPU

openai fields:

Field Notes
base_url Default http://localhost:1234/v1
model Model name sent in the request
stream Default true; false disables TTFT measurement
extra_body Merged into the request body for server-specific options

ovgenai fields:

Field Values Notes
model_path path OpenVINO IR directory, not a GGUF
device CPU | GPU | NPU
pipeline llm | vlm VLMs need vlm even for text-only prompts
attention_backend PA | SDPA Paged Attention requires OpenVINO 2026.2+

Paged Attention stores the KV cache in fixed-size blocks rather than one growing contiguous allocation per sequence, which keeps per-token latency flat as context grows. If throughput is far below published figures, check this first.

Two things that cause problems

Thinking mode. OVMS auto-detects a reasoning parser for Qwen3.6. The model then emits chain-of-thought into a separate reasoning_content field and leaves content empty until reasoning completes, which under any modest token cap means never. Every case scores zero. Disable it per target:

"extra_body": { "chat_template_kwargs": { "enable_thinking": false } }

The harness detects this case and raises a diagnostic naming the fix rather than reporting 0% accuracy. Whether thinking helps is task-dependent, and targets.example.json includes a thinking variant so you can measure it; run that one with --max-tokens 1024 or it truncates.

Constrained decoding and the # character. With a JSON schema active, generation terminates when the model attempts to emit # inside a string value. Output is cut mid-value and finish_reason is reported as stop. See Open issues for the full reproduction. This affects one of the ten bundled cases and is why the benchmark reports 90% rather than 100%.

Output

target          field acc   exact   json     TTFT    decode t/s  prefill t/s
q36-ovms-streamed   90.0%   90.0%    90%    2.89s      20.5 ± 1           62
  • field acc: individual fields correct. Primary metric.
  • exact: cases with every field correct. Predicts unattended use.
  • json: responses that parsed.
  • TTFT: time to first token, dominated by prefill. Blank when not streaming.
  • decode t/s: median after the first token, ± standard deviation.
  • prefill t/s: prompt tokens ÷ TTFT.

A per-field table follows, which is usually where the actionable information is. A model failing only on dates is a prompt problem, not a capability problem.

Results

Core Ultra 7 258V (Lunar Lake), 32 GB LPDDR5X, Arc 140V iGPU, OVMS 2026.3, OpenVINO 2026.3. 10 invoice-extraction cases, 3 repeats.

Non-streamed Streamed
Field accuracy 90.0% ¹ 90.0% ¹
Exact case (all 6 fields) 90.0% ¹ 90.0% ¹
Decode n/a ² 20.5 tok/s ± 1
Prefill n/a ² 62 tok/s
TTFT (median) n/a ² 2.89 s

¹ The same single case fails in both modes, and it is the constrained-decoding # bug described in Open issues, not a wrong answer. The model returns every field correctly up to the point generation is terminated. Nine of ten cases are exact.

² Non-streaming cannot measure time-to-first-token, so prefill and decode cannot be separated. Use the streamed target for performance, the non-streamed target for a second accuracy sample.

Intel's 42-43 tok/s on a 368H is a different tool, serving path and runtime version, so the two are not directly comparable. See Why. Decode is bandwidth-bound and Lunar Lake's on-package LPDDR5X runs ~136 GB/s, which would account for some of the gap.

Prefill is compute-bound and integrated graphics are weak there. The prefill figure is prompt tokens divided by TTFT at roughly 190-token prompts, so it includes queueing, tokenisation and the first decode step. It is a lower bound, not isolated prefill throughput. A sweep across prompt lengths is needed to separate the two.

Results from other hardware are welcome. See Contributing.

Console

moe-on-arc Console

A local web console for driving the model and the harness. Standard library only, no build step, binds to loopback.

python ui/server.py                 # then open http://localhost:7860
Tab What it does
Chat Streams responses from the model, with TTFT, token count and tok/s under each one. Drop an image in to use the vision side of the model. Thinking mode is off by default and toggleable.
Benchmark Renders results.json: accuracy per target, per-field grid, throughput, failing cases.
Run Picks a task, targets file and label, starts the harness and streams its progress live. Writes results.json and refreshes the dashboard.

Options: --ovms for the endpoint (default http://localhost:8000/v3), --model, --port.

The console only ever launches moe_on_arc.py, and task and target filenames are validated against the project root, so a stray request cannot run an arbitrary command. One harness run at a time.

Open issues

Constrained decoding terminates on # in a string value. With response_format set to a JSON schema, generation stops when the model attempts to emit # inside a string. The value is cut mid-string and finish_reason is reported as stop, so the failure is not visible from the response metadata alone.

probe_hash.py reproduces it. Standard library only, imports no project code, and prints the exact request body with --show-request. It exits 1 on reproduction.

python probe_hash.py --repeat 5

Measured on OVMS 2026.3, OpenVINO 2026.3, GPU device, OpenVINO/Qwen3.6-35B-A3B-int4-ov, temperature: 0, thinking disabled:

Input With schema Without schema
INV-#-99 truncated 5/5, emitted # 0/5 truncated 0/5, emitted # 5/5
INV-%-99 truncated 0/5 truncated 0/5
INV-@-99 truncated 0/5 truncated 0/5
INV-&-99 truncated 0/5 truncated 0/5
INV-X-99 truncated 0/5 truncated 0/5

Controls with other punctuation pass 20/20 under a schema, and the same input completes normally with no schema, so this is specific to # on the constrained-decoding path. Not tested on --target_device CPU.

Reported upstream: openvinotoolkit/model_server#4439

Other issues:

  • OpenVINO GenAI has no schema-constrained decoding, so ovgenai targets generate unconstrained JSON and are not like-for-like against constrained targets. Use --no-schema everywhere for a fair comparison.
  • --no-schema is not a workaround for the # bug. Without the schema the model renames fields (supplier_name, grand_total) and sometimes does not return JSON at all, so constrained decoding is doing real work.
  • NPU targets are limited to small models. Larger ones exceed the plugin's per-graph memory budget, sometimes failing silently.
  • Energy is not measured. Joules-per-token often decides laptop deployments and needs external instrumentation.
  • Prefill throughput is not isolated. TTFT at short prompts is dominated by fixed per-request overhead. Measuring across prompt lengths would separate prefill from that overhead.

Troubleshooting

Symptom Cause
0% accuracy, empty responses Thinking mode. Add extra_body with enable_thinking: false.
One case truncates mid-value The # bug. See Open issues.
CL_OUT_OF_RESOURCES Shared GPU memory budget. See setup step 1.
Option not found: MAX_PROMPT_LEN --max_prompt_len is NPU-only. Drop it.
/v3/models returns empty Still loading. REST binds before the model is ready.
Cannot parse JSON body from curl PowerShell escaping. Use a request file with -Encoding ascii.
DefaultCPUAllocator: not enough memory Running a model conversion. The downloaded artifact is already converted.
Throughput far below published Check attention_backend; PA vs SDPA matters at this scale.
No GPU enumerated Update the Intel graphics driver directly from Intel.
First run appears to hang Graph compilation. Cached afterwards, excluded from timings.
Import errors after installing OpenVINO Transformers version.

Contributing

Results from other hardware are the most useful contribution. Include:

  • Full check_env.py output
  • The exact targets.json entry used
  • Repeat count and the complete output table including variance
  • Model artifact identifier: a Hugging Face repo

A benchmark is a runtime version, a model artifact, a device, a prompt shape and an attention path. Results without that context are not comparable.

probe_hash.py output from other devices and OVMS versions is also useful, particularly CPU.

Contents

File Purpose
moe_on_arc.py The harness
check_env.py Pre-flight environment verification
ui/server.py Local console: chat, dashboard, run control
ui/index.html Console front end, single file
probe_hash.py Standalone reproduction for the # bug
task_invoices.json Bundled task: 10 invoice-extraction cases
targets.example.json Working target definitions. Copy to targets.json
requirements.txt Pinned dependencies

License

Apache License 2.0. See LICENSE and NOTICE.

Adam Milton-Barker is an Intel Software Innovator. This project is independent work and is not an Intel product. Views and results here are the author's own and have not been reviewed or endorsed by Intel Corporation or Alibaba Group. OpenVINO is a trademark of Intel Corporation; Qwen is a trademark of Alibaba Group.

About

Benchmarking Qwen3.6 MoE on Intel integrated graphics. Measuring accuracy, not just throughput.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages