Search
Advertisement
OpenAI introduces CriticGPT, an AI tool that helps coders identify bugs and improve code quality

OpenAI introduces CriticGPT, an AI tool that helps coders identify bugs and improve code quality

OpenAI has unveiled CriticGPT, an AI model designed to enhance code reviews by identifying errors in code generated by ChatGPT. In trials, CriticGPT improved code review outcomes by 60% compared to traditional methods.

Danny D'Cruze
Danny D'Cruze
  • New Delhi,
  • Updated Jun 28, 2024 9:34 AM IST
OpenAI introduces CriticGPT, an AI tool that helps coders identify bugs and improve code qualityOpenAI launched GPT 4o (Photo-Unsplash)

OpenAI has introduced CriticGPT, a new AI model based on GPT-4, designed to identify errors in code produced by ChatGPT. In trials, CriticGPT improved code review outcomes by 60% when used compared to those who did not.

CriticGPT is set to be integrated into OpenAI's Reinforcement Learning from Human Feedback (RLHF) labeling pipeline, aiming to provide AI trainers with better tools to evaluate complex AI outputs.

The GPT-4 models that power ChatGPT are designed to be helpful and interactive through RLHF. This process involves AI trainers comparing different responses and rating their quality. As ChatGPT's reasoning improves, its mistakes become subtler, making it harder for trainers to identify inaccuracies. This highlights a key limitation of RLHF: advanced models can become so knowledgeable that human trainers struggle to provide meaningful feedback.

CriticGPT has been trained to write critiques that highlight inaccuracies in ChatGPT's answers. Although its suggestions are not always perfect, they significantly help trainers identify more issues than when working without AI assistance. In experiments, teams using CriticGPT produced more comprehensive critiques and identified fewer false positives compared to those working alone. A second trainer preferred the critiques from the Human+CriticGPT team over those from an unassisted reviewer more than 60% of the time.

CriticGPT was trained using a method similar to ChatGPT but focused on identifying mistakes. AI trainers inserted errors into ChatGPT's code and provided example feedback. These trainers then compared multiple critiques of the modified code to evaluate CriticGPT's performance. CriticGPT's critiques were preferred in 63% of cases involving naturally occurring bugs, partly because it produced fewer unhelpful complaints and fewer hallucinated problems.

Despite its success, CriticGPT has limitations. It was trained on short ChatGPT answers and needs further development to handle longer, more complex tasks. Additionally, while models still hallucinate and trainers occasionally make labeling mistakes, the focus on single-point errors needs to expand to address errors spread across multiple parts of an answer.

Advertisement

Related Articles

For Unparalleled coverage of India's Businesses and Economy – Subscribe to Business Today Magazine

ABOUT THE AUTHOR

Danny D'Cruze
Danny D'Cruze

When Danny is not tinkering with his latest gadget you'll find him in the kitchen whipping up a new recipe. To balance out all the eating, he enjoys working out and taking long walks.

When he's not immersed in his hobbies, Danny is a social butterfly who loves meeting new people and getting to know them over a good conversation. He's also an avid gamer who can spend hours lost in a virtual world, and a dance enthusiast who's ever ready to hit the dance floor.

Danny loves long drives and podcasts. He's always eager to share his latest discovery and can talk for hours about his favorite shows. 

Published on: Jun 28, 2024 9:34 AM IST