From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality
Abstract
Code review helps maintain software quality before code integration, but it also imposes a substantial workload on human reviewers. As generative artificial intelligence becomes part of software development, code review is shifting from a primarily human review process toward AI-supported review processes in which large language model (LLM) reviewers and AI agent reviewers participate alongside human reviewers. However, we still lack empirical evidence on how this transition affects review efficiency and review quality. In this paper, we study 1.02 million reviewed pull requests from 207 GitHub projects that transition across three code review eras: human-centric review, LLM-assisted review, and agentic code review. We identify three AI reviewer adoption practices: Gradual AI Adoption, Rapid LLM Adoption, and Rapid AI Agent Adoption. We further model pull request review discussions as reviewer interaction sequences to characterize how human, LLM, and AI agent reviewers collaborate during the review process. Our results show that agent-involved collaboration patterns, especially reviews initiated by AI agents or involving multiple AI agents, are associated with faster review decisions under Gradual AI Adoption and Rapid AI Agent Adoption. However, these efficiency gains do not translate into better review quality. We also find that review activity and pull request type remain important across eras, while human-AI collaboration patterns become the strongest explanatory factor for review efficiency once LLM and AI agent reviewers participate. These findings provide empirical guidance for designing AI-supported code review processes that improve efficiency without weakening review quality.
Community
We study 1.02 million reviewed pull requests from 207 GitHub projects to examine how code review evolves across human-centric, LLM-assisted, and agentic review eras. By modeling review discussions as human–AI interaction sequences, we find that AI-agent involvement is associated with faster review decisions, especially when agents initiate reviews or multiple agents participate. However, these efficiency gains do not consistently improve review quality. Our findings highlight both the promise and limitations of agentic code review and provide practical guidance for designing human–AI review workflows.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Early Adoption of Agentic Coding Tools by GitHub Projects (2026)
- Predicting Acceptance and Review Effort in Human and Agent Pull Requests (2026)
- Is Agentic Code Review Helpful? Mining Developers'Feedback to CodeRabbit Reviews in the Wild (2026)
- Empirical Study on the Characteristics and Evolution of AI-usage in GitHub Repositories: Evidence from Code Comments (2026)
- Understanding the Rejection of Fixes Generated by Agentic Pull Requests -- Insights from the AIDev Dataset (2026)
- Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents (2026)
- Augmentation with Dilution: A Large-Scale Empirical Study of Human Contributor Ecosystems After AI Coding Agent Adoption (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.13196 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper