{"id":1457,"date":"2023-11-05T23:25:00","date_gmt":"2023-11-05T20:25:00","guid":{"rendered":"https:\/\/www.bengusuozcan.com\/?p=1457"},"modified":"2025-05-12T22:46:31","modified_gmt":"2025-05-12T19:46:31","slug":"metaculus-agi-predictions","status":"publish","type":"post","link":"https:\/\/www.bengusuozcan.com\/en\/future\/metaculus-agi-predictions\/","title":{"rendered":"Metaculus AGI Predictions"},"content":{"rendered":"<p>Mar 2025 Edit: <br>&#8211; Did a recent update on the numbers. <br>&#8211; I guess I&#8217;ve also updated what I consider AGI &#8211; not assuming that people fully buy in this Metaculus prediction, but sharing mine if I were to anchor on this. I feel like novel scientific research needs to be part of the picture: an AI model building and testing a hypothesis across many different scientific fields, including those that require physical research components and orchestration of research groups, and being able to produce reputable scientific research or file a patent over a discovery. <br>&#8211; Also, the Ferrari assembly doesn&#8217;t seem to fit for sufficient multimodality. There\u2019s a clear business incentive for visual processes in car assembly to be supported by existing data, but that doesn\u2019t necessarily speak to broader generalization. We need more everyday, real-world scenarios to really test a model\u2019s generalizability. This isn\u2019t the best example, and I wouldn\u2019t call it AGI if a model could reason through it, but it\u2019s in the direction I\u2019m thinking: your toaster is broken, you&#8217;re leaving for work in 10 minutes, and all your pans are in the dishwasher. Does your helper agent consider that the best move might be to quickly wash a pan and toast the bread on the stove\u2014rather than trying to fix the toaster or start a whole new meal from scratch?<\/p>\n\n\n\n<p>August 2024 Edit: This was originally posted in November 2023. I keep the text almost fully to its original, only update the lines about the main prediction if I&#8217;ve significantly updated my prediction.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<p>I feel so lost in conversations these days. So surprised that it\u2019s just been a year since ChatGPT is out, I feel like it\u2019s been here all my lifetime. Especially after moving out, so much of my conversations is going on in online chats, and it\u2019s been difficult ot communicate. I don\u2019t care abour my timelines at all, but this post is going to save me a tone of time so here it is \ud83d\ude00&nbsp;<\/p>\n\n\n\n<p>Metaculus AGI prediction:<\/p>\n\n\n\n<p>I think the tasks here are widely different from each other, combining purely cognitive and purely physical-cognitive tasks together, leading to a clear bottleneck, but anyway.<\/p>\n\n\n\n<p><strong>Q&amp;A Benchmark: 2025-2026<\/strong><\/p>\n\n\n\n<p>MMLU, pretty similar benchmark, is already at &gt;80%, so this should be solved in 2025 or 2026.<\/p>\n\n\n\n<p><strong>APPS (Algorithmic Reasoning): 2027<\/strong><\/p>\n\n\n\n<p>MATH and SWE-bench are kind of indicative for this benchmark both? And they seem to have doubled per year, so APPS should catch up? 2027 feels right, idk.<\/p>\n\n\n\n<p><strong>Turing Test (2-hour, multi-modal): ~2030<\/strong><\/p>\n\n\n\n<p>This one\u2019s tricky. Depends on the judge, this convo can take an entire day or just 10 mins. I guess what matters here is the context length and multi-modality. Also, does AI need to be embedded in a hardware system that has continuous sensory input? Idk, I&#8217;ll index this similar to the Ferrari assembly question below, anchoring on difficulty of processing sensory information.<\/p>\n\n\n\n<p><strong>Ferrari Assembly (Real Robots!): 2035-2037<\/strong><\/p>\n\n\n\n<p>Are we talking about Ferrari actually implementing this because if they have &#8220;some&#8221; business understanding, the answer is &#8220;probably never&#8221; because they should keep the luxury human-hand-made brand. I mean we still pay a tone for concerts even though we have Spotify, so someone will need handmade cars, right? <\/p>\n\n\n\n<p>Whatever, in terms of ability, I still think that specific tasks require mastery, like AI will probably produce a car that is functional, but it may need a call-back because a small detail was not perfect. So I&#8217;d anchor this on robotics and it has a long way to go. I think it&#8217;s plausible that AI can instruct someone to make the car by like 2030, but it&#8217;ll probably take another 5 years until it can make it, so maybe by 2035?<\/p>\n\n\n\n<p><strong>Big Picture Prediction: <\/strong>2035, anchored on the Ferrari task.<\/p>\n\n\n\n<p><strong>Range?<\/strong> 2032-2038. +\/- 3 years. Why? Idk.<\/p>\n\n\n\n<p><strong>Assumptions<\/strong>: <\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Robotics: there might be a breakthrough, like data farms work out? idk.<\/li>\n\n\n\n<li>Compute? Will scale &#8211; at least to support on the cognitive tasks.<\/li>\n\n\n\n<li>Regulations: won&#8217;t slow down much<\/li>\n\n\n\n<li>Warning shots: hopefully not by then, but bio stuff is scary<\/li>\n<\/ul>","protected":false},"excerpt":{"rendered":"<p>Mar 2025 Edit: &#8211; Did a recent update on the numbers. &#8211; I guess I&#8217;ve also updated what I consider AGI &#8211; not assuming that people fully buy in this Metaculus prediction, but sharing mine if I were to anchor on this. I feel like novel scientific research needs to be part of the picture: an AI model building and testing a hypothesis across many different scientific fields, including those that require physical research components and orchestration of research groups, and being able to produce reputable scientific research or file a patent over a discovery. &#8211; Also, the Ferrari assembly doesn&#8217;t seem to fit for sufficient multimodality. There\u2019s a clear business incentive for visual processes in car assembly to be supported by existing data, but that doesn\u2019t necessarily speak to broader generalization. We need more everyday, real-world scenarios to really test a model\u2019s generalizability. This isn\u2019t the best example, and I wouldn\u2019t call it AGI if a model could reason through it, but it\u2019s in the direction I\u2019m thinking: your toaster is broken, you&#8217;re leaving for work in 10 minutes, and all your pans are in the dishwasher. Does your helper agent consider that the best move might be to [&hellip;]<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_stc_notifier_status":"sent","_stc_notifier_sent_time":"2025-04-26 23:46:09","_stc_notifier_request":false,"_stc_notifier_prevent":false,"_stc_subscriber_keywords":"","_stc_subscriber_search_areas":"","footnotes":""},"categories":[90],"tags":[],"class_list":["post-1457","post","type-post","status-publish","format-standard","hentry","category-future"],"_links":{"self":[{"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/posts\/1457"}],"collection":[{"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/comments?post=1457"}],"version-history":[{"count":4,"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/posts\/1457\/revisions"}],"predecessor-version":[{"id":1465,"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/posts\/1457\/revisions\/1465"}],"wp:attachment":[{"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/media?parent=1457"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/categories?post=1457"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.bengusuozcan.com\/en\/wp-json\/wp\/v2\/tags?post=1457"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}