{"id":15445,"date":"2025-08-08T03:50:40","date_gmt":"2025-08-08T03:50:40","guid":{"rendered":"https:\/\/demo4.dedicatedhost247.com\/newstime\/openai-gets-caught-vibe-graphing\/"},"modified":"2025-08-08T03:50:40","modified_gmt":"2025-08-08T03:50:40","slug":"openai-gets-caught-vibe-graphing","status":"publish","type":"post","link":"https:\/\/demo4.dedicatedhost247.com\/newstime\/openai-gets-caught-vibe-graphing\/","title":{"rendered":"OpenAI gets caught vibe graphing"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1ymtmqpi _17nnmdy1 _17nnmdy0 _1xwtict1\">During its <a href=\"https:\/\/www.theverge.com\/openai\/748017\/gpt-5-chatgpt-openai-release\">big GPT-5 livestream on Thursday<\/a>, OpenAI showed off a few charts that made the model seem quite impressive \u2014 but if you look closely, some graphs were a little bit off.<\/p>\n<\/div>\n<div>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1ymtmqpi _17nnmdy1 _17nnmdy0 _1xwtict1\">In one, ironically showing how well GPT-5 does in \u201cdeception evals across models,\u201d the scale is all over the place. For \u201ccoding deception,\u201d for example, the chart shown onstage says GPT-5 with thinking apparently gets a 50.0 percent deception rate, but that\u2019s compared to OpenAI\u2019s smaller 47.4 percent o3 score which somehow has a larger bar. OpenAI appears to have accurate numbers for this chart in its <a href=\"http:\/\/I added that he says it\u2019s correct in the blog post\u2026\">GPT-5 blog post<\/a>, however, where GPT-5\u2019s deception rate is labeled as 16.5 percent.<\/p>\n<\/div>\n<div>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1ymtmqpi _17nnmdy1 _17nnmdy0 _1xwtict1\">With <a href=\"https:\/\/x.com\/EgeErdil2\/status\/1953505551570415718\">this chart<\/a>, OpenAI showed onstage that one of GPT-5\u2019s scores is <em>lower<\/em> than o3\u2019s but is shown with a bigger bar. In this same chart, o3 and GPT-4o\u2019s scores are different but shown with equally-sized bars. It was bad enough that CEO Sam Altman commented on it, <a href=\"https:\/\/x.com\/sama\/status\/1953513280594751495\">calling it<\/a> a \u201cmega chart screwup,\u201d though he noted that a correct version <a href=\"https:\/\/www.theverge.com\/openai\/748017\/gpt-5-chatgpt-openai-release\">is in OpenAI\u2019s blog post<\/a>.<\/p>\n<\/div>\n<div>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1ymtmqpi _17nnmdy1 _17nnmdy0 _1xwtict1\">An OpenAI marketing staffer also <a href=\"https:\/\/x.com\/pranaveight\/status\/1953517360071299113\">apologized<\/a>, saying, \u201cWe fixed the chart in the blog guys, apologies for the unintentional chart crime.\u201d<\/p>\n<\/div>\n<div>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1ymtmqpi _17nnmdy1 _17nnmdy0 _1xwtict1\">OpenAI didn\u2019t immediately respond to a request for comment. And while it\u2019s unclear if OpenAI used <a href=\"https:\/\/www.theverge.com\/openai\/748017\/gpt-5-chatgpt-openai-release\">GPT-5<\/a> to actually make the charts, it\u2019s still not a great look for the company on its big launch day \u2014 especially when it is touting the \u201csignificant advances in reducing hallucinations\u201d with its new model.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/www.theverge.com\/news\/756444\/openai-gpt-5-vibe-graphing-chart-crime\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>During its big GPT-5 livestream on Thursday, OpenAI showed off a few charts that made the model seem quite impressive<\/p>\n","protected":false},"author":1,"featured_media":15446,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[2],"tags":[],"_links":{"self":[{"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/posts\/15445"}],"collection":[{"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/comments?post=15445"}],"version-history":[{"count":0,"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/posts\/15445\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/media\/15446"}],"wp:attachment":[{"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/media?parent=15445"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/categories?post=15445"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/demo4.dedicatedhost247.com\/newstime\/wp-json\/wp\/v2\/tags?post=15445"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}