TLDR overview
- GPT-6.1 Sol passes 85.29% of the 544 HumanEval and MBPP tasks with executable tests, against 83.09% for GPT-6 Sol.
- GPT-6.1 Sol generated 669,155 lines of code against 480,125 for GPT-6 Sol, which is 39.4% more for the same benchmark.
- Every density measure improved. Bug density fell 3.3%, vulnerability density 32.4%, code smell density 5.2% and overall issue density 5.5%.
- Total findings rose 31.7%, from 9,859 to 12,982
- Vulnerabilities are the one place both readings agree. The count fell from 137 to 129 across 39% more code.
Compared with GPT-6 Sol, GPT-6.1 Sol writes 39% more code, lifts its pass rate by 2.20 points, and produces fewer findings per line on every measure we track.
GPT-6.1 Sol writes more code than either GPT-6 variant, recovers most of the pass rate gap with GPT-6 Astra, and gets cleaner per line in the process.
We ran it through Sonar's LLM evaluation framework against the same Java benchmark we use for every model, and against the same GPT-6 Sol run we published in September.
How was GPT-6.1 Sol evaluated?
Model: GPT-6.1 Sol, medium reasoning effort
Baseline: GPT-6 Sol, medium reasoning effort
Language: Java
Benchmark: the same Java benchmark we use for every model, covering HumanEval, MBPP and ComplexCodeEval. Both runs record 4,444 tasks. All of them are analyzed for code quality and security, which is where every density figure below comes from. 544 of them ship with executable tests, and the pass rate covers only those.
Analyzer: SonarQube systematic code analysis. Complexity and code smell densities are per 1,000 lines of code (kLOC). Bug and vulnerability densities, and all category breakdowns, are per million lines (mLOC).
Three terms worth defining first:
- Cyclomatic complexity: counts independent paths through a function.
- Cognitive complexity: a SonarQube metric that weights nested and deeply branched logic more heavily, reflecting how hard the code is for a human to read.
- Severity: the impact level a finding carries, inherited from the rule that raised it. It describes how much a finding affects the software qualities Sonar measures. It is not a prediction of how likely that code is to reach production.
How do GPT-6.1 Sol and GPT-6 Sol compare on key metrics?
All four finding density rows came down, both complexity measures stayed roughly the same, the pass rate went up and missing completions halved. Lines of code is the row that went the other way, up 39.4%, and it is the one the rest have to be read against.
What is GPT-6.1 Sol's functional pass rate?
85.29% across the 544 test-backed tasks, against 83.09% for GPT-6 Sol. That is 2.20 points higher.
Across the Sol line the pass rate has climbed with each release. GPT-5.6 Sol passed 81.99%, GPT-6 Sol passed 83.09%, and GPT-6.1 Sol passes 85.29%. That makes it the strongest of the Sol variants we have measured, and it lands just under GPT-6 Astra's 85.85%.
How much code does GPT-6.1 Sol generate compared to GPT-6 Sol?
GPT-6.1 Sol generated 669,155 lines across the benchmark. GPT-6 Sol generated 480,125. That is 39.4% more code for the same 4,444 tasks.
It used 67,363 functions to do that, up from 50,077, a 34.5% increase. Classes grew more slowly, from 17,737 to 19,943, which is 12.4%. So each class carries more: about 34 lines against 27.
For scale, 669,155 lines puts GPT-6.1 Sol slightly above GPT-6 Astra, which generated 656,445. On volume this release looks much more like Astra than like the variant it replaces.
Comment lines moved the furthest of anything in the run. Comment density is 3.3% of output, up from 1.3%, which in absolute terms is 23,126 comment lines against 6,219. That is roughly 3.7 times as many comments.
Anyone maintaining this code later gets considerably more inline context than GPT-6 Sol provided.
Does GPT-6.1 Sol generate more complex code than GPT-6 Sol?
Both complexity measures remained approximately the same.
Cyclomatic complexity is 223.34 per kLOC, down from 225.56, a 1.0% decrease. Cognitive complexity is 164.71 per kLOC, down from 167.38, a 1.6% decrease.
Neither is a meaningful move on its own at that size. The useful reading is the direction set against the volume. Logic is marginally less dense per line, but it is spread across 39% more lines, so the total amount of branching and nesting in the output went up substantially even though the per-line figures edged down.
What is GPT-6.1 Sol's bug density compared to GPT-6 Sol?
Bug density is 608 per mLOC, down from 629. A 3.3% decrease.
Blocker bug density fell 29%, from 56 to 40 per mLOC, while both runs produced exactly 27 blocker bugs. Read that as no change rather than an improvement. The same 27 findings are spread across 39% more code, so anyone triaging blocker bugs has identical work either way. It is the clearest case in this run of a density figure that looks like progress on its own and is not.
Critical is the row that moved against GPT-6.1 Sol, and it moved on both readings. The rate rose 57%, from 44 to 69 per mLOC, and the count rose from 21 findings to 46. Major came down 15% by rate and minor rose 7%.
The density improvement is real but small, and the volume increase outweighs it. GPT-6.1 Sol produced 407 bugs across the benchmark against 302, which is 105 more to work through. So the 3.3% better rate per line does not mean less bug triage. Each line is marginally less likely to carry a bug, while the stack of bugs to clear is about a third taller.
Concurrency and threading is the largest bug category in both runs, and here it is the one that got worse on both measures. The rate rose 21%, from 242 to 293 per mLOC, and the count rose 69%, from 116 findings to 196. It now accounts for 48% of all bugs in the output, up from 38%.
Exception handling moved the same way, up 67% by rate and from 16 findings to 37.
Most other categories improved by rate. Null and data value halved, control flow fell 76%, resource leaks fell 34%, performance and structure fell 25%. Read the counts alongside those. Null and data value went from 21 findings to 15 and control flow from 8 to 3, so these are small absolute movements that look larger as rates because the denominator grew. Performance and structure actually rose slightly in count, from 54 findings to 56, while its rate fell 25%.
Concurrency findings are the expensive kind. They are hard to reproduce, they depend on the environment they run in, and they surface as intermittent failures rather than clean ones. The rules that flag them are narrow and specific: double-checked locking, a lock not released on every path out of a method, synchronizing on a field that gets reassigned, a counter incremented on a volatile field, and calls to wait or notify made without holding the lock. Those turn up in singleton and cache setup, connection and worker pools, shared counters, and producer and consumer queues.
With 196 of them in the output, this is the first place to point review effort.
What security vulnerabilities does GPT-6.1 Sol generate in code?
This is the strongest result in the run, and the only area where both readings point the same way.
Vulnerability density is 193 per mLOC, down from 285. A 32.4% decrease. The absolute count came down too, from 137 findings to 129, which is 5.8% fewer across 39% more code.
Blocker is the one security row that rose, and the numbers behind it are very small. Blocker vulnerability density went from 2 to 4 per mLOC, which is 1 finding in GPT-6 Sol's output and 3 in GPT-6.1 Sol's. Three blocker vulnerabilities across 669,155 lines is still a clean result at the top of the scale. At that size a single finding moves the rate by more than a point, so the rate is the wrong way to read this row.
Critical remains the dominant tier at 75% of security density, against 69% for GPT-6 Sol.
The top three categories all move the same way, and it is a denominator story. Each produced about the same number of findings in both runs: cryptography misconfiguration 50 against 51, insecure system resource handling 44 against 41, and injection attack 18 in both. Their rates fell 29%, 22% and 27% because the code around them grew by 39%, not because fewer of them appeared.
The genuine reduction is inadequate error handling, down from 8 findings to 2.
Hard-coded credentials is a new finding type in this run. GPT-6 Sol produced none and GPT-6.1 Sol produced 2, at 3 per mLOC.
Cryptography misconfiguration and insecure system resource handling together account for 141 of the 193 per mLOC, so roughly three quarters of the security surface sits in two categories. In counts that is 94 of 129 findings. That has been true of every model in this family, and it is the answer to where a pipeline should point first. Crypto misconfiguration covers weak algorithms, insecure key sizes, and random number generators used in ways they should not be.
How maintainable is GPT-6.1 Sol generated code?
Code smell density is 18.60 per kLOC, down from 19.62. A 5.2% decrease.
In absolute terms smells rose, from 9,420 findings to 12,446, which is 32.1% more.
The mix moved in a useful direction. The three highest tiers all came down by rate: blocker 23%, critical 7% and major 16%. The two lowest rose, minor by 14% and info by 9%.
In counts every tier rose, which is what 39% more code produces. Blocker went from 25 findings to 27, critical from 1,553 to 2,022, major from 4,949 to 5,818, minor from 2,526 to 4,022 and info from 367 to 557. Minor is where most of the growth sits.
Collection and generics stays the largest category by a wide margin at 43% of all smells, with its rate down 7%. One caveat on reading that row, and it applies to every row in the table. Each category groups a lot of rules, seventeen Java rules in this case, and they are not all the same kind of problem. This one spans raw types and missing type parameters, redundant casts, and collection handling that could be simpler, so the total does not tell you how much of it is any one of those.
Deprecated API usage is the only named category that improved on both readings, down 38% by rate and from 142 findings to 122. That reverses what we flagged as the one to watch in the GPT-6 Sol evaluation.
How does GPT-6.1 Sol code volume affect total findings?
Total findings rose 31.7%. Bugs rose 34.8% and code smells 32.1%. Vulnerabilities fell 5.8%.
Every density measure went the other way. Bug density down 3.3%, vulnerability density down 32.4%, code smell density down 5.2%, overall issue density down 5.5%.
The gap between those two readings is the 39.4% volume increase. Findings grew by 31.7% while code grew by 39.4%, and that difference is the whole of the density improvement.
Three denominators, two answers:
- Per line of code, GPT-6.1 Sol is cleaner. 19.40 findings per kLOC against 20.53.
- Per task, it produces more. 2.92 findings per task against 2.22. The task count is fixed at 4,444 in both runs, so that is another way of saying 12,982 findings against 9,859.
For a team sizing total review effort, GPT-6 Sol produced less work. For a team asking how carefully each line needs reading, GPT-6.1 Sol is the cleaner output. And for a team weighing correctness, GPT-6.1 Sol passes 12 more of the 544 test-backed tasks.
Vulnerabilities are the exception worth pausing on, because they are the one category where both readings agree. GPT-6.1 Sol produced fewer of them in absolute terms across considerably more code. That is not a denominator effect, and it is the clearest single improvement in the run.
How do GPT-6.1 Sol and GPT-6 Sol compare on token usage?
Input tokens were identical across both runs at 1,322,069, which is what you would expect from models using the same tokenizer.
Output tokens rose 4.1%, from 6.30 million to 6.56 million. Set against 39.4% more code, that means GPT-6.1 Sol spent considerably fewer tokens per line of output. Per line of code it used 9.81 output tokens against 13.13, which is 25.3% fewer.
Reasoning tokens are where the two diverge, and they diverge in the opposite direction from the last release. GPT-6.1 Sol used 1.28 million against GPT-6 Sol's 2.64 million, so it cut its reasoning budget by 51.4% while producing more output. Reasoning fell from 41.8% of output tokens to 19.5%.
Total tokens came out 3.4% higher for GPT-6.1 Sol, at 7.88 million against 7.63 million.
Of the two, GPT-6 Sol was the one that spent more on thinking and less on writing. GPT-6.1 Sol reverses that, and it has the higher pass rate of the pair.
What are the biggest risks of using GPT-6.1 Sol generated code in production?
GPT-6.1 Sol is the more correct and the cleaner of the two per line, and it produces more work overall because it writes more code. Three places to point verification effort.
Concurrency is first. It is the largest bug category by a wide margin, and it rose on both readings. The rate rose 21% and the count rose 69%, from 116 findings to 196, which is 48% of all bugs in the output and more than three times the next category. These are among the harder findings to catch in review and to reproduce once they ship. The rules that flag them are narrow and well covered by automated analysis, which is the practical answer here. Exception handling rose on both readings too, up 67% by rate and from 16 findings to 37, so it belongs in the same pass even though it stays much smaller.
The critical reliability tier is second. It rose 57% by rate and from 21 findings to 46. Blocker bugs held at exactly 27 findings in both runs, so this is a critical-tier movement rather than a shift toward the most severe findings. If your quality gate weights critical findings heavily, that is the row that will show up.
Review volume is third. 12,982 findings against 9,859 is about 3,100 more items, and 2.92 findings per task against 2.22. Per line the code is cleaner, but there is more of it, so the pile of work to triage is larger. That is a planning question rather than a quality one, and it is the honest cost of the extra volume.
Security is the good news, and it is unambiguous. Fewer vulnerabilities in absolute terms across 39% more code, density down 32.4%, and the critical tier flat in count while the codebase around it grew. Blocker findings are 3 across 669,155 lines. Cryptography misconfiguration and insecure system resource handling still account for about three quarters of the security surface, and both are well covered by automated analysis.
Three takeaways:
- More code, cleaner per line. 39.4% more code than GPT-6 Sol for the same tasks, with every density measure down and the pass rate up 2.20 points to 85.29%, the strongest of the Sol variants we have measured. The trade is 12,982 findings to work through against 9,859.
- Security improved on both readings. 129 vulnerabilities against 137, across 39% more code, with density down 32.4%. This is the one finding type where volume growth did not mean more to fix.
- Concurrency is the biggest bug category. 196 findings, which is 48% of all bugs and more than three times the next category. It rose on both readings, up 21% per line and up 69% in count, so it is the first place to point review effort.
Full evaluation results for GPT-6.1 Sol, along with every other model we have measured, are on the Sonar LLM Leaderboard.

