Three months on LUMI: AI often has the right answer. It just cannot pick it.
Three months and 4,703 GPU hours on a Finnish supercomputer. When model collaboration helps, what fine-tuning ThinkingCap achieved, and why choosing the right answer can be harder than generating it. Results, charts and datasets to investigate yourself.

My project Fusion Playground on the Finnish supercomputer LUMI is nearing completion. In three months, I used 4,703 of the allotted 5,000 GPU hours, ran over twenty experiments, and collected over 305,000 scored answers. I do almost all the research and development myself, with support from the Czech LUMI AI Factory team at IT4Innovations.
I arrived in July with a simple question: When are multiple AI models working together better than one strong model? And how much does it cost?
The most interesting finding turned out to be about something other than model size. On difficult science questions, at least one of Gemma’s eight attempts was correct 95.5% of the time. Yet voting selected a correct answer only about 86% of the time. The solution was often already there. It just got lost on the way to the user.
What I take away from LUMI
- Generating more answers is not enough. In this experiment, the largest untapped opportunity was selecting the right one.
- A capable model with concise reasoning is hard to beat. Gemma-4-31B scored 91.3% in the main quality comparison and 91.4% on the new set. The panels and routers I tested did not surpass it on average.
- Collaboration makes sense where we have a good reason for it. Tests helped with programming, repeated attempts helped with some math problems. However, this did not result in a universal recipe.

The same eight attempts per question, using Gemma-4-31B on 198 GPQA questions. The left figure requires access to the answer key. It is not the accuracy of a system I can deploy today.
Eight answers from AI. Which one should you trust?
Imagine a meeting in which someone says the right solution, but the majority favors another. This is exactly what happened in my experiments. More generation doesn't mean better decisions.

Solid lines: actual voting result. Dashed lines: at least one correct answer among the attempts. This upper bound uses knowledge of the solution. Even 32 attempts did not raise Gemma-4-31B’s voting accuracy above 86.4%.
I call it the selector wall. Reliably picking the right answer from those eight candidates could add almost ten percentage points without a new base model. It would not be free: generating eight answers and checking them both have a cost.
Letting another AI decide is an obvious next step. In a separate IFEval instruction-following experiment, however, the judges I tested disagreed with rule-based checks in 20 to 25% of cases. When asked again, they changed their verdict on the same answer in 9 to 25% of cases. This does not rule out AI judges. It means these particular judges and settings need validation. The figures also cannot be assumed to apply to GPQA science questions.
Checks with a firm basis worked best: run the program’s tests, recompute the result, check the formatting rules. A model saying “this looks right” is a different kind of evidence.
What does model collaboration actually mean?
In Hyperspace I develop HyperFusion, a system for collaboration between models, and Atlas, which selects the appropriate procedure based on the task. To read the results, it is useful to distinguish four things:
| Procedure | What's going on |
|---|---|
| Voting | The model answers repeatedly, or several models vote. The most common answer wins. |
| Selection | Choose one existing candidate, for example using test results. |
| Cascade | A cheaper model answers first. If its answer cannot be adequately verified or fails a check, a stronger model takes over. |
| Synthesis | Another model combines the candidates into a new answer. |
The results of one procedure are not automatically the results of the others. In the final experiment, I measured actual synthesis on instruction-following tasks and questions over documents. Other tasks were mainly voting, selection or cascade. For the code, public tests were used for selection and separate hidden tests for final evaluation.
Hyperspace also works with large hosted models that have different prices and features. Therefore, I do not draw a universal verdict on its production fusion from the measurements of open models on LUMI.
One strong model is hard to beat
The final comparison included eleven model configurations and fifteen types of tasks. Alongside science questions and mathematics, these included programming, text processing, instruction following and Czech school-leaving exams. There were usually four to eight trials per task in the main matrix; some follow-up tests used more repetitions.
A router selects a model or method based on the task. I compared the original rules and learned routing policies with a simple strategy: always use the same model.

Top left is better. The reference choice A → B knows the results of other attempts on the same questions, so it is not a deployable router. Optimistic selection based on the evaluated results alone is not shown here.
| Procedure | Main set | New types of tasks | Price main / new* |
|---|---|---|---|
| Qwen3.6-35B-A3B without reasoning | 85.2% | 88.7% | 1.17 / 0.71 |
| Qwen3.8-27B with reasoning | 89.9% | 89.9% | 4.91 / 3.44 |
| Gemma-4-31B with reasoning | 91.3% | 91.4% | 2.01 / 1.62 |
| Panel of three models | 89.1% | 88.5% | 3.56 / 1.30 |
| Cascade | 90.7% | 90.1% | 3.96 / 1.96 |
| Learned routers | 90.2 to 91.3% | 90.3 to 91.4% | according to the setting |
| Reference selection A → B | 92.1% | 92.8% | 1.51 / 0.96 |
Thousandths of a GPU hour per task. This is an average score on a specific mixture of tasks, not a universal ranking of the models. A GPU hour here means an hour of work for one AMD MI250X module.
Gemma combined strong performance with concise reasoning. In the main set, the cheaper Qwen used about 1/1.7 of the GPU time, but was six points behind. The router thus had little room for error. With a different pair of models and different prices, the situation can be significantly more favorable.
How much confidence should we place in the numbers?
The main set contained 2,544 unique problems, of which 2,537 entered the strict comparison with all necessary answers. The new set added 1,455 problems of six additional types. The router was locked before I opened those results. Policy selection and evaluation also used separate repetitions, and training the routers involved splits by unique question as well.
Not every tenth of a point is an improvement. For example, 91.3% for Gemma has a 95% confidence interval of approximately 90.3 to 92.3% in the main analysis. Therefore, I describe small differences as observed results, not as proven superiority. Repeated answers to the same question are also not independent new questions.
Later experiments with other recipes and verifiers reused data whose results had already been inspected. They are useful for finding the next direction, but need another separate test. This is especially true for savings with JEV below.
When did multiple attempts really pay off?
| Situation | What worked | Observed result |
|---|---|---|
| Programming with Tests | Cheap model, verification, then a stronger model if needed | Comparable quality in about half the GPU time: 85.2% for 3.0 vs. 84.8% for 5.6 thousandths of a GPU hour |
| Programming with an emphasis on quality | Repeat strong model and check tests | Approximately +1.4 to +1.9 points for a 17 to 40% higher price |
| Competition mathematics | Voting multiple attempts or appropriate panel | In one comparison +3.4 points for 2.2 times the price |
| Specialist science questions | Different models or cascade | Around +1.3 to +2.0 points, for two to six times the cost |
| Search and calculation from text | Cheap model without extended reasoning | Similar or better quality, three to seven times cheaper |
In mathematics, I was surprised by how much the pattern of errors mattered. The smaller Gemma, with reasoning disabled, improved considerably with repeated attempts. The stronger Gemma-4-31B barely improved. Voting helps when attempts differ in useful ways. Repeating the same mistake does not make it disappear in a majority vote.

AIME contains 60 competition problems. A solid line is a vote, a dashed line means at least one correct attempt. The small number of questions is another reason not to overestimate minor differences.
ThinkingCap: better math in 30 GPU hours
One of the most encouraging results came from fine-tuning ThinkingCap-Qwen3.6-27B from BottleCap AI, Tomáš Mikolov's company. The model already aims to keep its reasoning concise. I wanted to see if I could push it further.
For the variant shown in the graph, I used 1,400 mathematics examples and 469 science and programming answers generated by the original model. For simpler mathematics, I chose short correct solutions, for more difficult examples, longer ones. The original answers were to help preserve other abilities. One such training session took approximately 30 GPU hours. The whole series of trials and evaluations cost more.

MATH-500: 500 questions, three times each, i.e. 1500 attempts per variant. The same fine-tuned variant gained 4.1 percentage points and shortened reasoning by 42%.
More accurate and at the same time more concise than the original ThinkingCap. For a specific math set, this is a very nice result for a small training budget.
But it comes at a price: the same variant dropped 3.3 points on AIME competition mathematics and 5.6 points on GPQA Science questions. Another of the five recipes lost less on science, but it was a different model. Their best numbers cannot be combined into one winner. I improved a specific math ability, didn't beat ThinkingCap in everything.
Before training, I checked both exact and near text overlaps with tests MATH-500, AIME, and GPQA; the check used found none. It is not a complete guarantee that similar tasks did not occur in the original pre-training. And it's a version based on Qwen3.6: BottleCap meanwhile released a newer ThinkingCap-Qwen3.8 against which I'm not measuring this result.
A better verifier instead of a bigger model?
Atlas uses the fast specialized model JEV from TypeSafe to estimate input properties. I also tried it to check answers that had already been generated. In these follow-up trials, it performed better than estimating the difficulty of the question alone.
| A signal for deciding when a cheap model is enough | AUC | Quality at about half the GPU time of Gemma |
|---|---|---|
| Original JEV questions about task properties | 0.77 | 88.8% |
| New JEV questions about difficulty | 0.79 | 89.3% |
| Text properties without JEV | 0.61 | 88.3% |
| Answer check with Gemma-4-31B | 0.77 | 89.4% |
| Checking the answer with Qwen3.6 | 0.82 | 88.8% |
| Checking answer with JEV | 0.84 | 89.8% |
AUC describes how well the signal ranks cases according to risk of error. A value of 0.5 corresponds to random order, 1 to perfect. An AUC of 0.84 does not mean 84% accuracy of answers.
In this setting, saving approximately 30% of GPU time cost half a percentage point of quality. Saving half the GPU time cost about one and a half points. These are points selected from a measured curve on already used data, not a ready-made guarantee for new customer inquiries. API calls to JEV cost an additional $0.00005 per task; that dollar cost is not included in the GPU-hour figures.
JEV is still a model that can make mistakes. It is not the equivalent of an executable test. I also tried the other verifiers in a specific economical setting, not in all possible forms. Nevertheless, it is a more promising guide for further work than a mere impression of the difficulty of the assignment.
A new generation can matter more than billions of extra parameters
Even the models themselves changed within three months. Compared with the equally sized Qwen3.6-27B, Qwen3.8-27B gained roughly 13 percentage points on competition mathematics and 18 on programming in these tests. At the same time, it consumed approximately half of the GPU time on these tasks.

Reasoning was enabled for both models. The improvement was not universal: on some tasks with text and Czech mathematics, the newer model lost slightly. It makes sense to continuously measure the models on your own tasks.
The opposite surprise came from MiMo-V2.6-Pro with approximately one trillion parameters, the largest model in the final comparison. It fit on one node in a compressed form with an average of 3.5 bits per parameter.
| MiMo-V2.6-Pro, 3.5 bit | Gemma-4-31B | |
|---|---|---|
| Science Questions GPQA | 78.3% | 85.9% |
| Czech language school-leaving exam | 92.2% | 91.0% |
| GPU hours per GPQA question | approximately 0.52 | approximately 0.005 |
MiMo did well on the Czech language exam, but was much more expensive and less accurate on GPQA. A fifth of the difficult questions hit the response-length limit. The higher-precision MXFP4 variant scored 4.7 points higher on the 64 questions completed by both variants, with no truncated answers. That is an interesting clue, but also a small, selected sample.
The reason may be compression, the quality of a particular version of the model, the way it is launched and the limits set. MiMo ran via llama.cpp, Gemma via vLLM. In addition, MiMo uses the MoE architecture, which activates only part of the parameters for each token. This is not a pure "trillion vs. 31 billion" experiment and cannot be used to condemn the model as such.
I interpret the earlier paid-API comparisons with the same caution:
| A small comparison test | Via API | At LUMI |
|---|---|---|
| GPQA, 138 questions | Gemini 3.1 Pro 93.5%, Claude Sonnet 5 86.2% | Gemma-4-31B 84.1% |
| GLM-5.2, scored open-ended task | 60 points out of 60 | 4-bit version: 35 points out of 60 |
| Kimi K3, 40 programming tasks | 31 correct | 2-bit version: 28 correct |
The last line means about 90% of the relative result in this test, not keeping 90% of all Kimi's abilities. The differences between the API and the local version also do not in themselves isolate the effect of quantization, i.e. compression of the numerical weights of the model. This may work well, but each specific file and execution method needs to be checked.
Two findings that would be a shame to forget
Long documents will test a different ability than the short benchmark. In the document question test, I added distracting text and extended the context from about 8 to 120k tokens. Tested Qwens lost about 2-5 points F1 score, gpt-oss about 18. This is not a verdict on all long documents, but it shows well why context-window size is not enough as a measure of quality.
Who does the judging matters. In an earlier attempt at summarization, the synthesis came out the winner in 92.7% of the comparisons when evaluated by the same model that produced it. With the independent evaluator Llama, it was 57.3%. The evaluator has changed, so it is not an accurate measurement of a single cause. But the difference is enough as a warning against the model issuing a report card to itself.
How this relates to other research
My experiments are not the first work on the cooperation of models. They help me put general ideas into specific costs and constraints on LUMI:
- Self-Consistency shows the benefit of multiple paths to solving and voting. My results remind me that it also depends on how similar mistakes the trials make.
- Mixture-of-Agents explores the synthesis of multiple model responses. This is a different process from voting itself.
- RouteLLM solves the choice of model according to the ratio of quality and price. The result of such a router always depends on which models it chooses.
- Let's Verify Step by Step trains to check the individual steps of the solution. My data usually only evaluates the final answer, so for a similar approach, it would be necessary to add another kind of annotations.
- Quantization Inflates Reasoning describes that compression can make reasoning longer. It's a possible explanation for part of my observations, not proof of causation in a particular MiMo run.
What I learned about the supercomputer itself
During the final experiments, up to 140 scheduler jobs ran concurrently. Each handled many queries. This concurrency allowed approximately 2,900 GPU hours to fit into one day and night; it does not mean that one user waited so long for a response.
The division of labor made a big difference. One LUMI node has four MI250X modules, a total of eight compute dies (GCDs). For a model that fits two chips, I could run four separate copies instead of one copy across eight chips.

Qwen3.6-35B-A3B: around 1015 vs 322 tokens per second. I'm measuring the aggregate throughput of a node over many queries, not the speed of a single response. Smaller copies will help if the model fits in memory and they have enough to work with.
I also continuously read the GPU power. According to those measurements, the final experiments consumed 562 kWh for accelerators. It does not include the entire server, network or cooling.

Energy per token helps compare serving efficiency. But for a real application, energy for a correctly solved task is more important. A shorter answer can be more economical overall even with a more expensive token.
What went wrong and how much it cost
Some results cost an unnecessary amount of time. Twice I set the thinking limit too short. Regenerating truncated answers consumed 777 GPU hours, roughly a sixth of the budget. I also learned not to trust interim scores: easier questions finish first, making the early results look better.
In the first analysis, I also compared panels from one trial with individual models averaged over multiple trials. I fixed the methodology; the main table above already uses separate repetitions and comparable ratings.
For the coding agent on SWE-bench, the success rate went from 42 to 61%, but the credit was due to better tools and work management, not retraining. Even a negative result saves another dead end.
| Item | Result |
|---|---|
| Allocation used | 4,703 of 5,000 GPU hours, reserve 297 |
| Final Experiments | approximately 2,900 GPU hours |
| Energy of the final | 562 kWh, GPU only |
| Graded outputs | 305,051 answers and another 5,314 program solutions |
| Paid API | about $42, of which about $0.85 for JEV in total |
305k answers is not 305k different questions. Number includes different models and repetitions. The GPU hours listed are the consumption of the grant allocation, not the bill I paid in cash.
Three months that I want to build on

I described the first experiences in the articles about LUMI and Kajaani heating and about first fusions and GLM-5.2. In July, LUMI AI Factory also wrote about the project.
In September, I spoke about experiments at CERNA.AI in Ostrava's Gong, a festival with more than two thousand participants. The lecture resulted in an article 20 + 5 = 29. Verified. and LUMI AI Factory wrote again about the event and cooperation. In parallel, I started a separate quantum project at the Czech VLQ; its results are not part of the experiments reported here on LUMI.

I see the next research step in better answer selection and verification. I have an extensive set of graded outputs to start with. But first, I need to separate the data for learning and the really new test so that the checker doesn't just remember familiar questions. I want to release the data, scripts and frozen sets for independent verification.
I have also started working with Radim Vavřík to prepare a POP performance analysis of my AI workloads on LUMI. I have supplied the test case and supporting material; the analysis itself has yet to begin. The aim is to find where computation stalls unnecessarily and how to use the hardware more effectively.
For Hyperspace, I take away a simple working rule: start with a good stand-alone model and add more complex collaborations where I measure the real benefit. So the most interesting question for the next months is: How can we get the right answer to the user when the AI has already generated it?
Another allocation, Aitta and an education model
My first experience with LUMI has been very positive. I am therefore considering applying for a larger follow-up allocation of computing time. A natural next step could be Fast Lane Access, which allows applications for up to 50,000 GPU hours over a maximum of three months. This is a plan, not an allocation that has already been awarded.
I definitely want to try Aitta, CSC’s service for running open-weight models on LUMI through an OpenAI-compatible API. It handles model setup and requests resources from the scheduler. For research and prototyping, it could save me substantial infrastructure work. Its documentation does, however, note that availability for production use is not guaranteed.
Three directions look particularly worthwhile:
- A verifier that can identify the right answer. Test on new questions whether a smaller fine-tuned model can select candidates better than voting, and whether it is worthwhile after including the verifier’s own cost.
- A model for education. Start by fine-tuning an open model to identify errors in a student’s reasoning and offer useful hints instead of simply revealing the answer. I want to compare it with the base model and with a model that retrieves supporting material. Answer accuracy alone is not enough; the quality of the guidance matters too.
- Noetica as a cognitive architecture. Build on my earlier Noetica experiments by connecting memory, planning, Atlas, model collaboration and result verification. I would measure each component’s contribution by disabling it in turn while keeping the compute budget constant.
These directions fit together: an educational task needs a good explanation, a verifier checks its correctness, and the architecture decides when a simple method is enough. I would use a larger budget to test that collaboration on new tasks, rather than just produce another table of bigger models.
Acknowledgments and where to start your own project
I thank EuroHPC JU, LUMI AI Factory, IT4Innovations, CSC (IT Center for Science) and the LUMI consortium for computing time, support and machine operation. Thanks also to the organizers CERNA.AI for the invitation.
I personally thank Jakub Siwek, who was the first to approach me and introduce me to Filip Oborník. He subsequently invited me to his podcast.

From left Jan Tyl, Jakub Siwek and Filip Oborník.
We acknowledge EuroHPC JU for awarding the project ID EHPC-AIF-2026PG01-843 access to LUMI at CSC, Finland.
Have your own idea? Contact the Czech team LUMI AI Factory in IT4Innovations at ai-factory@it4i.cz. I started with a smaller project myself.
And if you want to try the cooperation of models directly, enter Hyperspace.
Datasets and materials for replication
Here are the sources of fifteen types of tasks from the final comparison. Counts indicate unique questions, not the number of responses from all models. Links lead to source datasets; I haven't always used their entire current release.
Main set: 2,544 tasks
| Dataset | Part used | Count |
|---|---|---|
| GPQA | Diamond, previously divided into 60 and 138 questions, the same ordering of answer options | 198 |
| MATH-500 | Full test | 500 |
| AIME 2024 and AIME 2025 | Frozen sets, 30 problems from each year | 60 |
| IFEval | 541 entries, strict evaluation of all instructions in response | 541 |
| HotpotQA | Custom selection from the distractor variant | 120 |
| LiveCodeBench | Frozen selection of LCB-60 from file test6 | 60 |
| CERMAT: Czech | Multiple-choice questions | 649 |
| CERMAT: Mathematics | Multiple-choice questions | 126 |
| MMLU in Czech | Five questions per field loaded, seed 42; the saved selection has 290 items | 290 |
2,537 tasks entered the strict router comparison: seven IFEval items lacked required outputs. In this main table, HotpotQA is scored by exact answer match; for the long context experiment, I present F1. These are not interchangeable metrics.
New task types: set A8, 1,455 questions
| Dataset | Part used | Count |
|---|---|---|
| BIG-Bench Hard | 15 items from each of the 27 configurations | 405 |
| MMLU-Pro | Test, 25 items from 14 fields | 350 |
| Belebele | Test, Czech version ces_Latn | 200 |
| CzechBench Agree | Grammatical agreement test | 150 |
| DROP | Selection from the validation set | 200 |
| Klokan QA | Test, configuration balanced | 150 |
The A8 selection script uses seed 42. Dataset loading order matters too, so the seed alone does not fully specify the selection.
Download list of datasets, selected task IDs and checksums (JSON)
The file contains the identifiers of all 3,999 tasks across both sets and SHA-256 hashes of the saved input files. It also lists generation settings from the script and examples of runtime versions from the logs. It does not redistribute questions or answer keys; you can get them from the authors of the datasets on their terms. For example, GPQA requires acceptance of access terms.
For an exact replication of the whole experiment, I still need to publish the complete package: prompts, exact model and dataset revisions, corrected limits for individual runs, scoring scripts and the locked router policies. This list is a first concrete basis, not a claim that the entire experiment is already reproducible with a single command.
I fine-tuned ThinkingCap on selected examples from OpenR1-Math-220k and the original model’s own responses to prompts from Mixture-of-Thoughts. These are training sources, separate from the evaluation sets above.