A comparison of two AI code-review models reveals significant trade-offs between cost and accuracy. According to testing by Entelligence, GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens, while GPT-6 Astra costs $10 and $50 per million tokens respectively.

On 50 public benchmark pull requests from Cal.com, Sentry, Discourse, Keycloak, and Grafana, Luna found 69 verified bugs at a total cost of $0.20, compared to Astra’s 92 verified bugs costing $5.66. Luna achieved this at 23 seconds per review versus Astra’s 36 seconds, generating 3.1 times more output tokens while remaining far cheaper due to its lower output pricing.
However, precision differs substantially. Luna raised 93 findings with 74% precision, meaning about one in four comments was incorrect. Astra raised 96 findings with 96% precision, with only 4 findings failing verification. According to the analysis, developers accustomed to skimming AI review comments “will skim harder when a quarter of them are noise.”
Performance varied significantly by codebase and bug type. In Sentry, Discourse, and Grafana, Luna came within two verified bugs of Astra. In Cal.com, the gap widened to 21 versus 30 bugs. Keycloak, an identity and access management system, showed the largest disparity: Luna found 6 verified bugs to Astra’s 14, with only 50% of Luna’s Keycloak findings verified versus 93% for Astra.
On security bugs specifically, Luna found 9 of 24 versus Astra’s 19. Examples of security issues Astra caught that Luna missed include federated recovery codes never marked as used (allowing reuse) and global view permissions overriding individual client denials. Luna found bugs Astra missed in other categories: 25 bugs were identified only by Luna, mostly data and logic bugs.
The researchers note Luna is “good enough for everyday correctness bugs at that price” but recommend against using it alone for authentication or permission code review. Running both models on every pull request would identify 117 of 143 verified bugs (82%) for $5.86 total, adding 25 more verified bugs beyond Astra alone.
The testing methodology used the same prompt on identical diffs, with both GPT-6 Astra and GPT-5.6 Sol serving as judges to verify findings. All pull requests predate both models’ training cutoffs, and models received only diffs without repository history or production data.
Key facts
- Luna found 69 verified bugs at $0.20 total cost; Astra found 92 at $5.66
- Luna has 74% precision vs. Astra’s 96% precision on findings
- Luna completed reviews in 23 seconds vs. Astra’s 36 seconds
- On security bugs, Luna found 9 of 24 vs. Astra’s 19 of 24
- Performance gap widest on Keycloak (identity management code): Luna 6 bugs vs. Astra 14
