Code quality is measurable across several dimensions: reliability bugs, maintainability issues, security findings, readability, testing, and verification. Recent studies show that maintainability problems dominate large codebases, while AI-assisted development is increasing both perceived quality and the need for systematic checking.
Contents
- How often quality issues appear
- Language-level issue density
- Reliability and maintainability rates
- Security weaknesses and remediation
- What controlled Copilot testing measured
- What developers report about AI tools
- The AI verification gap
How often quality issues appear
Sonar analyzed code during the last six months of 2024 and identified approximately 445 million total issues. The figure is a broad count across the analyzed codebase, so it should be read as an observed volume of findings rather than a universal defect rate. The underlying source is The State of Code, Volume 1: Reliability — June 2025 and related Sonar reports.
Maintainability findings made up the largest category. About 420 million of the approximately 445 million issues were categorized as maintainability issues or code smells, according to The State of Code, Volume 3: Maintainability — July 2025. That is about 16 times the combined count of reliability bugs, vulnerabilities, and security hotspots in Sonar’s dataset. The result suggests that code quality work is often less about one dramatic failure and more about accumulated complexity, duplication, and difficult-to-change code.
Reliability findings were smaller in absolute volume but still substantial: about 16 million issues were categorized as reliability bugs. Sonar also categorized approximately 1.3 million identified issues as security vulnerabilities and about 8.4 million as security hotspots. These categories are not interchangeable. A vulnerability is a security weakness, while a hotspot is a security-sensitive issue requiring review; neither should automatically be treated as a confirmed exploit.
Sonar’s aggregate density benchmarks provide another perspective. Reliability findings averaged about 2,100 bugs per million lines of code analyzed, while security findings averaged roughly 170 vulnerabilities per million lines. The findings come from the last six months of 2024 and describe Sonar’s analyzed code, not every software project or every language.
Language-level issue density
Sonar’s language comparison shows why a single “average code quality” number can mislead. Issue counts and issue densities varied by language and by the amount of code represented in the dataset.
| Language | Issues or issue density in Sonar’s dataset | Measurement period |
|---|---|---|
| Java | About 106 million issues; 69,000 per million lines | Last six months of 2024 |
| JavaScript | About 116 million issues; 93,000 per million lines | Last six months of 2024 |
| TypeScript | About 62 million issues; 37,000 per million lines | Last six months of 2024 |
| Python | About 18 million issues; 20,000 per million lines | Last six months of 2024 |
| C++ | About 20 million issues; 58,000 per million lines | Last six months of 2024 |
| PHP | About 21 million issues; 37,000 per million lines | Last six months of 2024 |
The Java portion contained about 106 million issues across roughly 1.5 billion lines of code. JavaScript contained about 116 million issues and averaged about 93,000 issues per million lines. TypeScript contained about 62 million issues and averaged about 37,000 per million lines. Python contained about 18 million issues and averaged about 20,000 per million lines.
C++ provides a useful comparison for readers working close to systems programming. Its portion contained about 20 million issues across roughly 350 million lines of code, averaging about 58,000 issues per million lines. C# averaged about 65,000 issues per million lines, while PHP averaged about 37,000 and contained about 21 million issues.
These are benchmark observations from The State of Code, Volume 4: Languages — July 2025. They are not a ranking of languages from “best” to “worst.” Project age, framework use, codebase size, rule configuration, and the kinds of systems represented can all affect the number of findings. A density benchmark is most useful when comparing similar projects under consistent measurement.
Reliability and maintainability rates
The same Sonar research translates aggregate findings into developer- and code-based rates. During the examined period, Sonar caught about 3 reliability issues per developer per month and about 72 code smells per developer per month. Those figures show how much quality tooling can surface during ordinary development, but they do not say that every finding became a production incident or required the same remediation effort.
On a lines-of-code basis, Sonar’s maintainability analysis averaged about 53,000 code smells per million lines of code. Its reliability analysis averaged about 2,100 bugs per million lines. The difference between these rates reinforces the broader pattern: maintainability findings were far more numerous than reliability findings.
For C programmers, the practical implication is that quality measurement should include both behavior and changeability. A program may pass its current tests while still accumulating code smells that make later fixes riskier. Reliability metrics help identify defects that can cause incorrect behavior; maintainability metrics identify friction that can make future correctness harder to preserve.
The source labels matter here. Reliability figures come from The State of Code, Volume 1: Reliability — June 2025, while maintainability figures come from The State of Code, Volume 3: Maintainability — July 2025. Both report observations from the last six months of 2024.
Security weaknesses and remediation
Sonar found roughly 1,100 security hotspots per million lines of code analyzed during the last six months of 2024. That density is much higher than the roughly 170 vulnerabilities per million lines reported in the same research, reflecting the fact that hotspots are review signals rather than confirmed vulnerabilities. The security measurements are reported in The State of Code, Volume 2: Security — July 2025.
GitLab’s April 2024 survey highlights process conditions around those findings. Sixty-seven percent of individual contributors said at least one-quarter of their working code came from open-source libraries, yet only 21% of organizations were using a software bill of materials to document software components. Fifty-two percent of security professionals said organizational red tape often slowed vulnerability fixes, and 55% said vulnerabilities were most commonly discovered after code was merged into a test environment.
The survey also found measurement uncertainty. Fifty-one percent of CxOs said their developer-productivity measurement methods were flawed or that they wanted to measure productivity but were unsure how. Forty-five percent were not measuring developer productivity against business outcomes. These are survey responses from GitLab Survey Reveals Tension Around AI, Security, and Developer Productivity within Organizations, not measurements of defect density.
Toolchain choices were also associated with different responses: 74% of respondents whose organizations used AI for software development wanted to consolidate their toolchain, compared with 57% of respondents whose organizations did not use AI. That comparison describes reported preferences, not a causal effect of AI adoption.
What controlled Copilot testing measured
GitHub’s randomized 2024 study, updated in February 2025, measured several dimensions of code quality among developers with and without GitHub Copilot access. Developers with Copilot had a 53.2% greater likelihood of passing all 10 unit tests. In the same controlled study, Copilot users averaged 18.2 lines of code per readability error, compared with 16.0 lines for developers without Copilot.
The Copilot group averaged 4.63 code readability errors, versus 5.35 for the non-Copilot group. GitHub reported that Copilot-authored code contained 13.6% more lines per readability error. Blind reviewers rated Copilot-authored code 3.62% higher for readability, 2.94% higher for reliability, 2.47% higher for maintainability, and 4.16% higher for conciseness. Developers were also 5% more likely to approve code authored with GitHub Copilot.
These results are controlled-study measurements, not a guarantee that AI assistance improves every codebase. They also cover particular tasks, participants, and evaluation conditions. The detailed source is Does GitHub Copilot improve code quality? Here’s what the data says. Passing unit tests and receiving higher blind-review scores are useful signals, but they do not replace project-specific tests, security analysis, or human review.
What developers report about AI tools
GitHub’s 2024 enterprise developer survey found that 90% of U.S. respondents reported improved code quality when using AI coding tools. The corresponding figures were 81% in India, 61% in Brazil, and 60% in Germany. These are reported perceptions, so they should not be treated as equivalent to controlled defect measurements.
Across the four surveyed countries, 60% to 71% of respondents said AI coding tools made adopting a new language or understanding an existing codebase easy. Between 23% and 29% said those tasks became very easy. More than 98% said their organizations had experimented with AI coding tools to generate test cases.
The distinction between “easy” and “very easy” is important: the data indicates broad perceived assistance, but it does not establish that the resulting code was secure, maintainable, or correct in production. The survey source is Survey: The AI wave continues to grow on software development teams.
The AI verification gap
Sonar’s 2026 survey describes a later stage of AI-assisted development. Seventy-two percent of developers who had tried AI said they used it every day, and AI-generated code accounted for 42% of all committed code reported in the survey. Developers expected that share to reach 65% by 2027; this is a forecast based on 2026 survey responses, not a measured future outcome.
Sixty-four percent of developers had started using autonomous AI agents. Yet developer toil remained nearly 24% of the work week regardless of AI-use frequency. The productivity story therefore includes persistent work that AI adoption had not removed in the survey.
Verification was the clearest warning sign. Ninety-six percent of developers said they did not fully trust AI-generated code to be functionally correct, but only 48% said they always checked AI-assisted code before committing it. In other words, the survey measured a gap between limited trust and consistent verification behavior.
These findings come from Sonar Data Reveals Critical “Verification Gap” in AI Coding: 96% Don’t Fully Trust Output, Yet Only 48% Verify It. They are 2026 survey results and should not be read as legacy findings independently verified here. For code quality work, the measured pattern supports keeping automated tests, static analysis, security checks, and human review in the commit path even when AI tools improve speed or readability scores.