By Adam Milton-Barker - CogniTech Systems Ltd. Intel Software Innovator.
moe-on-arc is a benchmark harness for measuring accuracy and throughput together when serving a local LLM, plus a verified end-to-end configuration for Qwen3.6-35B-A3B (35B parameters, 3B active, INT4) on an Intel Arc 140V iGPU via OpenVINO Model Server.
Intel published figures for this model on an AI PC: 42-43 tok/s on a Core Ultra 7 368H, measured with the LLM Bench tool from openvino.genai, in-process via VLMPipeline on OpenVINO 2026.2 (article).
That post is what convinced me this was worth trying. I had a 258V rather than a 368H, and I wanted the model as a service rather than a library, since anything I would build against it needs an HTTP endpoint. Serving through OVMS also adds continuous batching, prefix caching and scheduling that in-process inference does not have.
So the numbers here are not comparable to Intel's: different silicon, different serving path, different runtime version. Intel's post notes the same thing, that a benchmark is a runtime version, a model artifact, a device, a prompt shape and an attention path. This repo is one more configuration on that list, on hardware not covered by the published figures, with accuracy measured alongside throughput.
git clone https://github.com/CogniTechSystems/moe-on-arc
cd moe-on-arc
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python check_env.pycheck_env.py verifies versions, enumerates devices with their real memory properties, and checks headroom. Run it before downloading ~18 GB.
Qwen3.6 requires Transformers 5.2, which is why requirements.txt pins it. It is not pulled in by the OpenVINO or Optimum installs, so it must be installed explicitly.
Windows exposes about half of system RAM as the graphics budget, 16 GB on a 32 GB machine. The weights are 18.3 GB, so the default configuration fails with CL_OUT_OF_RESOURCES.
Intel Graphics Software → Graphics → General → Shared GPU Memory Override. Raise it, then confirm:
$ python check_env.py
--- versions ---
[ ok ] openvino 2026.3.0-22451-8a17657b995-releases/2026/3
[ ok ] openvino-genai 2026.3.0.0-3277-bd8d6542e3c
[ ok ] transformers 5.2.0
--- devices ---
enumerated: CPU, GPU, NPU
[ ok ] GPU: Intel(R) Arc(TM) 140V GPU (16GB) (iGPU)
(the size in that name is a static driver string, not the live budget)
total addressable 22.1 GB
raise via Intel Graphics Software > Graphics > General > Shared GPU Memory Override
[ ok ] NPU present (GPU is the recommended device for Qwen3.6-35B-A3B)
--- memory ---
[ ok ] system RAM 31.6 GB total
[warn] only 21.9 GB free - close things before loading
INT4 weights are ~19 GB and the iGPU shares this pool
--- disk ---
[ ok ] 484.9 GB free on this volume (model download needs ~20 GB)
Environment looks good. Next: download the model (see README.md).
22.1 GB against 18.3 GB of weights leaves room for the KV cache.
hf download OpenVINO/Qwen3.6-35B-A3B-int4-ov --local-dir ov_models/qwen3.6-35b-a3b-int4-ovThis artifact is already converted to OpenVINO IR and quantized to INT4. No conversion step is needed.
Download ovms_windows_<version>_python_off.zip from the model_server releases. The embedded interpreter is only needed for custom pipeline nodes.
$ovms = "C:\ovms\ovms\ovms.exe"
& $ovms --rest_port 8000 `
--model_path ov_models\qwen3.6-35b-a3b-int4-ov `
--model_name qwen3.6-35b-a3b `
--task text_generation `
--target_device GPU `
--tool_parser hermes3 `
--cache_dir .ov_cache `
--enable_prefix_caching true| Flag | Why |
|---|---|
--tool_parser hermes3 |
Parses tool calls out of model output. Required for agentic use. |
--cache_dir |
Caches the compiled graph. First load ~64 s, later loads much faster. |
--enable_prefix_caching |
Avoids re-prefilling a repeated system prompt. Valuable given weak prefill. |
This is the single-model form, so no config.json is needed; that file is only for serving several models, and --add_to_config requires a valid one to already exist.
Do not pass --max_prompt_len on GPU. It is NPU-only and fails with Option not found: MAX_PROMPT_LEN.
Wait for state changed to: AVAILABLE. The REST port binds before loading finishes, so an empty /v3/models early on is expected.
The OVMS endpoint path is /v3, not /v1. We will write the body to a file as PowerShell mangles inline JSON escaping, and -Encoding ascii avoids a BOM the parser rejects.
'{"model":"qwen3.6-35b-a3b","messages":[{"role":"user","content":"Reply with exactly: working"}],"max_tokens":20,"chat_template_kwargs":{"enable_thinking":false}}' |
Set-Content req.json -Encoding ascii
curl.exe http://localhost:8000/v3/chat/completions -H "Content-Type: application/json" -d "@req.json"Leave OVMS running in its window. In a second terminal, with the venv active:
copy targets.example.json targets.json
python moe_on_arc.py -t task_invoices.json -T targets.json --only q36-ovms -r 3 --max-tokens 512 --failures --save results.jsonpython moe_on_arc.py -t TASK -T TARGETS [options]
| Flag | Effect |
|---|---|
-t, --task |
Task JSON file (required) |
-T, --targets |
Targets JSON file (required) |
-r, --repeat N |
Runs per case; enables variance reporting |
--only LABEL |
Run a subset (repeatable) |
--no-schema |
Disable constrained decoding |
--failures |
Print failing cases |
--save FILE |
Write full results as JSON |
--max-tokens N |
Generation cap (default 256) |
--timeout SECONDS |
Per-request timeout (default 300) |
--quiet |
Suppress the progress counter |
A target is abandoned after two consecutive errors. A GPU OOM can leave the driver context in a state where further calls hang.
{
"name": "invoice extraction",
"system": "You extract structured data from invoice text...",
"template": "Extract the fields from this invoice:\n\n{text}",
"numeric_tolerance": 0.001,
"schema": { "type": "object", "properties": {}, "required": [] },
"cases": [ { "text": "raw input", "expected": { "field": "value" } } ]
}Every key in expected is graded independently. Numbers use numeric_tolerance (relative) and tolerate currency formatting; strings are compared case- and whitespace-insensitively; lists order-insensitively.
task_invoices.json holds ten invoice cases built to break things: a credit note with a negative total, a PO number adjacent to the invoice number, subtotals directly above grand totals, and dates in ISO, DD/MM/YYYY, Norwegian dotted and 31st January forms.
The harness is domain-agnostic. Swap in log lines, CVs, tickets, listings, anything with a known answer.
targets.json is a list of objects, each with a label and a backend.
| Backend | Transport | Works with |
|---|---|---|
openai |
/v1/chat/completions |
OVMS (on /v3), LM Studio, vLLM, llama-server, mlx_lm.server |
ovgenai |
in-process | OpenVINO GenAI on Intel CPU, iGPU, NPU |
openai fields:
| Field | Notes |
|---|---|
base_url |
Default http://localhost:1234/v1 |
model |
Model name sent in the request |
stream |
Default true; false disables TTFT measurement |
extra_body |
Merged into the request body for server-specific options |
ovgenai fields:
| Field | Values | Notes |
|---|---|---|
model_path |
path | OpenVINO IR directory, not a GGUF |
device |
CPU | GPU | NPU |
|
pipeline |
llm | vlm |
VLMs need vlm even for text-only prompts |
attention_backend |
PA | SDPA |
Paged Attention requires OpenVINO 2026.2+ |
Paged Attention stores the KV cache in fixed-size blocks rather than one growing contiguous allocation per sequence, which keeps per-token latency flat as context grows. If throughput is far below published figures, check this first.
Thinking mode. OVMS auto-detects a reasoning parser for Qwen3.6. The model then emits chain-of-thought into a separate reasoning_content field and leaves content empty until reasoning completes, which under any modest token cap means never. Every case scores zero. Disable it per target:
"extra_body": { "chat_template_kwargs": { "enable_thinking": false } }The harness detects this case and raises a diagnostic naming the fix rather than reporting 0% accuracy. Whether thinking helps is task-dependent, and targets.example.json includes a thinking variant so you can measure it; run that one with --max-tokens 1024 or it truncates.
Constrained decoding and the # character. With a JSON schema active, generation terminates when the model attempts to emit # inside a string value. Output is cut mid-value and finish_reason is reported as stop. See Open issues for the full reproduction. This affects one of the ten bundled cases and is why the benchmark reports 90% rather than 100%.
target field acc exact json TTFT decode t/s prefill t/s
q36-ovms-streamed 90.0% 90.0% 90% 2.89s 20.5 ± 1 62
- field acc: individual fields correct. Primary metric.
- exact: cases with every field correct. Predicts unattended use.
- json: responses that parsed.
- TTFT: time to first token, dominated by prefill. Blank when not streaming.
- decode t/s: median after the first token, ± standard deviation.
- prefill t/s: prompt tokens ÷ TTFT.
A per-field table follows, which is usually where the actionable information is. A model failing only on dates is a prompt problem, not a capability problem.
Core Ultra 7 258V (Lunar Lake), 32 GB LPDDR5X, Arc 140V iGPU, OVMS 2026.3, OpenVINO 2026.3. 10 invoice-extraction cases, 3 repeats.
| Non-streamed | Streamed | |
|---|---|---|
| Field accuracy | 90.0% ¹ | 90.0% ¹ |
| Exact case (all 6 fields) | 90.0% ¹ | 90.0% ¹ |
| Decode | n/a ² | 20.5 tok/s ± 1 |
| Prefill | n/a ² | 62 tok/s |
| TTFT (median) | n/a ² | 2.89 s |
¹ The same single case fails in both modes, and it is the constrained-decoding # bug described in Open issues, not a wrong answer. The model returns every field correctly up to the point generation is terminated. Nine of ten cases are exact.
² Non-streaming cannot measure time-to-first-token, so prefill and decode cannot be separated. Use the streamed target for performance, the non-streamed target for a second accuracy sample.
Intel's 42-43 tok/s on a 368H is a different tool, serving path and runtime version, so the two are not directly comparable. See Why. Decode is bandwidth-bound and Lunar Lake's on-package LPDDR5X runs ~136 GB/s, which would account for some of the gap.
Prefill is compute-bound and integrated graphics are weak there. The prefill figure is prompt tokens divided by TTFT at roughly 190-token prompts, so it includes queueing, tokenisation and the first decode step. It is a lower bound, not isolated prefill throughput. A sweep across prompt lengths is needed to separate the two.
Results from other hardware are welcome. See Contributing.
A local web console for driving the model and the harness. Standard library only, no build step, binds to loopback.
python ui/server.py # then open http://localhost:7860| Tab | What it does |
|---|---|
| Chat | Streams responses from the model, with TTFT, token count and tok/s under each one. Drop an image in to use the vision side of the model. Thinking mode is off by default and toggleable. |
| Benchmark | Renders results.json: accuracy per target, per-field grid, throughput, failing cases. |
| Run | Picks a task, targets file and label, starts the harness and streams its progress live. Writes results.json and refreshes the dashboard. |
Options: --ovms for the endpoint (default http://localhost:8000/v3), --model, --port.
The console only ever launches moe_on_arc.py, and task and target filenames are validated against the project root, so a stray request cannot run an arbitrary command. One harness run at a time.
Constrained decoding terminates on # in a string value. With response_format set to a JSON schema, generation stops when the model attempts to emit # inside a string. The value is cut mid-string and finish_reason is reported as stop, so the failure is not visible from the response metadata alone.
probe_hash.py reproduces it. Standard library only, imports no project code, and prints the exact request body with --show-request. It exits 1 on reproduction.
python probe_hash.py --repeat 5Measured on OVMS 2026.3, OpenVINO 2026.3, GPU device, OpenVINO/Qwen3.6-35B-A3B-int4-ov, temperature: 0, thinking disabled:
| Input | With schema | Without schema |
|---|---|---|
INV-#-99 |
truncated 5/5, emitted # 0/5 |
truncated 0/5, emitted # 5/5 |
INV-%-99 |
truncated 0/5 | truncated 0/5 |
INV-@-99 |
truncated 0/5 | truncated 0/5 |
INV-&-99 |
truncated 0/5 | truncated 0/5 |
INV-X-99 |
truncated 0/5 | truncated 0/5 |
Controls with other punctuation pass 20/20 under a schema, and the same input completes normally with no schema, so this is specific to # on the constrained-decoding path. Not tested on --target_device CPU.
Reported upstream: openvinotoolkit/model_server#4439
Other issues:
- OpenVINO GenAI has no schema-constrained decoding, so
ovgenaitargets generate unconstrained JSON and are not like-for-like against constrained targets. Use--no-schemaeverywhere for a fair comparison. --no-schemais not a workaround for the#bug. Without the schema the model renames fields (supplier_name,grand_total) and sometimes does not return JSON at all, so constrained decoding is doing real work.- NPU targets are limited to small models. Larger ones exceed the plugin's per-graph memory budget, sometimes failing silently.
- Energy is not measured. Joules-per-token often decides laptop deployments and needs external instrumentation.
- Prefill throughput is not isolated. TTFT at short prompts is dominated by fixed per-request overhead. Measuring across prompt lengths would separate prefill from that overhead.
| Symptom | Cause |
|---|---|
| 0% accuracy, empty responses | Thinking mode. Add extra_body with enable_thinking: false. |
| One case truncates mid-value | The # bug. See Open issues. |
CL_OUT_OF_RESOURCES |
Shared GPU memory budget. See setup step 1. |
Option not found: MAX_PROMPT_LEN |
--max_prompt_len is NPU-only. Drop it. |
/v3/models returns empty |
Still loading. REST binds before the model is ready. |
Cannot parse JSON body from curl |
PowerShell escaping. Use a request file with -Encoding ascii. |
DefaultCPUAllocator: not enough memory |
Running a model conversion. The downloaded artifact is already converted. |
| Throughput far below published | Check attention_backend; PA vs SDPA matters at this scale. |
| No GPU enumerated | Update the Intel graphics driver directly from Intel. |
| First run appears to hang | Graph compilation. Cached afterwards, excluded from timings. |
| Import errors after installing OpenVINO | Transformers version. |
Results from other hardware are the most useful contribution. Include:
- Full
check_env.pyoutput - The exact
targets.jsonentry used - Repeat count and the complete output table including variance
- Model artifact identifier: a Hugging Face repo
A benchmark is a runtime version, a model artifact, a device, a prompt shape and an attention path. Results without that context are not comparable.
probe_hash.py output from other devices and OVMS versions is also useful, particularly CPU.
| File | Purpose |
|---|---|
moe_on_arc.py |
The harness |
check_env.py |
Pre-flight environment verification |
ui/server.py |
Local console: chat, dashboard, run control |
ui/index.html |
Console front end, single file |
probe_hash.py |
Standalone reproduction for the # bug |
task_invoices.json |
Bundled task: 10 invoice-extraction cases |
targets.example.json |
Working target definitions. Copy to targets.json |
requirements.txt |
Pinned dependencies |
Apache License 2.0. See LICENSE and NOTICE.
Adam Milton-Barker is an Intel Software Innovator. This project is independent work and is not an Intel product. Views and results here are the author's own and have not been reviewed or endorsed by Intel Corporation or Alibaba Group. OpenVINO is a trademark of Intel Corporation; Qwen is a trademark of Alibaba Group.

