HOW TO

How to run Qwen3.8-Flash-Next on NVIDIA DGX Spark

Last updated:

Qwen3.8-Flash-Next is an AI model from Alibaba’s Qwen team that can answer questions, write and explain code, and understand images and video. Its model weights are available to download, so you can run it on your own hardware and connect it to chat apps or coding tools.

You'll understand how to set up a text chat on a 128GB NVIDIA DGX Spark, using llama.cpp to load the model and handle requests. You’ll download roughly 94GB of model files, build the server, and get your first response before connecting another app.

The Spark’s unified memory helps make room for a model this large: its CPU and GPU share the same 128GB pool.

What you’ll need

This setup uses a 128GB DGX Spark running NVIDIA DGX OS, with CUDA Toolkit 13.0 or later. You’ll need about 94GB of disk space for the model, additional space for the build tools, and internet access to GitHub and Hugging Face.

Open a terminal on the Spark, either directly or over SSH, and run:

Bash / Shell
nvidia-smi
export PATH=/usr/local/cuda/bin:$PATH
nvcc --version

nvidia-smi should identify the GB10. Use nvcc --version to find the installed CUDA Toolkit version; the CUDA number shown by nvidia-smi describes driver support.

The model files are public, so you don’t need a Hugging Face account or access token to download them. They’re released under the Qwen Community License 1.0, available on the model page.

The commands below follow the project documentation but have not yet been tested end to end on a Spark.

1. Download the model

We’re using Unsloth’s UD-IQ4_XS version. It stores the model in GGUF format, which llama.cpp can load, and uses quantization to reduce the space occupied by the weights.

Install the Hugging Face download tool in a Python virtual environment, then download the model:

Bash / Shell
sudo apt update
sudo apt install -y python3-venv

python3 -m venv ~/venvs/qwen-download
~/venvs/qwen-download/bin/python -m pip install --upgrade huggingface_hub

~/venvs/qwen-download/bin/hf download \
  unsloth/Qwen3.8-Flash-Next-GGUF \
  --revision efc50b7c3195a10eebd8d3e32f83b27b2aeb239d \
  --include "UD-IQ4_XS/*" \
  --local-dir ~/models/qwen38-flashnext-gguf

The download contains three files, known as shards, totaling approximately 93.7GB. They go into the UD-IQ4_XS folder inside ~/models/qwen38-flashnext-gguf. Keep all three together: you’ll point the server at the first file, and it will find the other two.

If the download stops before it finishes, run the same command again to resume it.

2. Build llama.cpp for the Spark

Next, build the program that will run the model. These commands follow NVIDIA’s llama.cpp build workflow, using version b11429, which supports Qwen3.8-Flash-Next.

Install the build tools:

Bash / Shell
sudo apt install -y \
  git cmake build-essential curl \
  libcurl4-openssl-dev libssl-dev

Then download llama.cpp and build its server with CUDA support for the GB10:

Bash / Shell
git clone --branch b11429 --depth 1 \
  https://github.com/ggml-org/llama.cpp \
  ~/llama.cpp-qwen38

cd ~/llama.cpp-qwen38

cmake -B build \
  -DCMAKE_BUILD_TYPE=Release \
  -DGGML_CUDA=ON \
  -DCMAKE_CUDA_ARCHITECTURES=121a-real

cmake --build build --config Release \
  --target llama-server -j 16

Once the build finishes, llama-server will be in ~/llama.cpp-qwen38/build/bin/.

3. Start the model server

The first run uses a 32,768-token context, which holds the prompt, conversation history, and generated response. It also uses one request slot, so the server handles one request at a time.

Run:

Bash / Shell
~/llama.cpp-qwen38/build/bin/llama-server \
  --model "$HOME/models/qwen38-flashnext-gguf/UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf" \
  --alias qwen38-flash-next \
  --host 127.0.0.1 --port 11434 \
  --n-gpu-layers 999 \
  --ctx-size 32768 \
  --batch-size 2048 \
  --parallel 1 \
  --jinja

The model path points to the first shard downloaded in Step 1. --n-gpu-layers 999 requests GPU offload for all available layers, while --jinja applies the model’s chat template to your messages. The name qwen38-flash-next is the alias you’ll use when connecting an app.

Keep this terminal open. Loading a model this size can take several minutes, and the startup messages will appear here. Ctrl+C stops the server.

Once the log says the model is loaded, open a second terminal on the Spark and run:

Bash / Shell
curl --fail-with-body -sS http://127.0.0.1:11434/health

The response {"status":"ok"} means it’s ready. If it reports that the model is still loading, wait and try again.

4. Get your first response

Start with a short question so you can see the request and answer together. In the second terminal, run:

Bash / Shell
curl --fail-with-body -sS \
  http://127.0.0.1:11434/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "qwen38-flash-next",
    "messages": [
      {"role": "user", "content": "What is 17 multiplied by 23? Reply with the number only."}
    ],
    "max_tokens": 128,
    "temperature": 0,
    "chat_template_kwargs": {"enable_thinking": false}
  }'

Look for 391 in the response. Thinking is disabled for this request so the model can answer directly within the short output limit.

Now ask for a longer response:

Bash / Shell
curl --fail-with-body -sS \
  http://127.0.0.1:11434/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "qwen38-flash-next",
    "messages": [
      {"role": "user", "content": "Count from 1 to 60, separated by spaces."}
    ],
    "max_tokens": 200,
    "temperature": 0,
    "chat_template_kwargs": {"enable_thinking": false}
  }'

Alongside the answer, llama.cpp returns a timings section. prompt_per_second describes how quickly it processed your input, while predicted_per_second measures how quickly it generated the reply.

You can also connect a chat app or coding tool from another computer. The server is bound to 127.0.0.1, so it accepts connections locally on the Spark. To reach it remotely, open an SSH tunnel from the computer running your app, replacing spark-user and spark-address with your login and the Spark’s address:

Bash / Shell
ssh -N -L 11434:127.0.0.1:11434 spark-user@spark-address

5. Give the model more context

A larger context lets you include more conversation history or longer documents in a request. To move from 32,768 to 65,536 tokens, stop the server with Ctrl+C, change --ctx-size 32768 to --ctx-size 65536, and run the launch command again. Use the health check before sending another request.

Context uses additional memory from the Spark’s shared pool. Run free -h to watch whole-system memory use, and return to the previous setting if the server runs out of memory. The model supports a native context of 262,144 tokens, but that doesn’t guarantee it will fit on this machine.

The model has to process the prompt before it starts answering, so a longer document can mean a longer wait for the first token. Repeated requests may reuse a cached prompt prefix, making them faster than the first request with new text.

How fast does it run?

On a 128GB Spark running UD-IQ4_XS with llama.cpp b11429 generated roughly 26–28 tokens per second. The counting response shown above reached 27.8 tokens per second, while the server log below shows approximately 25.6–25.8.

The server log separates the time spent processing the prompt from the time spent generating the answer.

Those figures cover the recorded requests, not long-context or multi-user benchmarks.

Memory use is substantial. During that earlier run, the system was using 103GiB out of 121GiB, with 4.2GiB of swap in use. The memory reported beside the GPU process accounts for only part of the system’s usage.

The system memory summary shows how much of the Spark’s shared pool is in use.

What about MTP and image input?

Multi-token prediction, or MTP, proposes several upcoming tokens for the main model to verify. Accepting several at once can speed up generation. Unsloth supplies a separate MTP head for this, but its shared head needs support that isn’t included in the build used here. See its MTP guide for the supported software.

Qwen3.8-Flash-Next also supports image and video input. Those inputs need a separate vision projector file, named mmproj, loaded alongside the model.

If something doesn’t work

Problem What to do
nvcc is not found Run export PATH=/usr/local/cuda/bin:$PATH, then try nvcc --version again.
The build rejects the GPU architecture Use CUDA Toolkit 13.0 or later and the 121a-real setting from Step 2.
unknown model architecture: 'qwen4exp' Run the executable built in Step 2; an older copy of llama-server may not support this model.
The model file cannot be opened Use the full Step 3 path, including UD-IQ4_XS, and finish downloading all three shards.
The server runs out of memory Reduce --ctx-size, stop other memory-intensive workloads, and restart.
Connection refused Look at the server terminal for a startup error. From another computer, keep the SSH tunnel open.
Empty answer with finish_reason: "length" Use the non-thinking request in Step 4. Requests with thinking enabled may need a larger output-token budget.

More room for local AI

CORSAIR PRO builds systems for AI inference, fine-tuning, and training, from compact and full-tower workstations to rackmount GPU servers and multi-node clusters. Our range includes workstation configurations with up to four double-width GPUs, with processor, graphics, system memory, and storage options tailored to the workload.

We can also prepare the software environment before delivery, installing, configuring, and testing the requested drivers, containers, and AI frameworks. We assemble and test each system against CORSAIR benchmarks and the workload it’s configured for. Our workstations include a three-year standard warranty, with on-site support options available.

The CORSAIR PRO configurator lets you explore available systems and select their components, with configured pricing as you make changes. You can review the complete build summary and request a quote for that configuration.

PRODUCTS IN ARTICLE

Stay up to date with CORSAIR. Get our latest News, Guides, and Product Updates in your Google feeds.

Add CORSAIR as a preferred source

JOIN OUR OFFICIAL CORSAIR COMMUNITIES

Join our official CORSAIR Communities! Whether you're new or old to PC Building, have questions about our products, or want to chat about the latest PC, tech, and gaming trends, our community is the place for you.

Was this article helpful?