Подтвердите e-mail

Для публикаций, комментариев, реакций и сообщений подтвердите адрес.

Профиль

Ethan Mollick

Профиль Vively

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so they hide all that stuff. They should instead explain choices like a good PM would (what should be delegated? what should be generalized? what tests to run?)

51158

Considered one of the best text adventure games (it is more interactive story, with few puzzles), 1985's A Mind Forever Voyaging is still worth a try. Since it is now open source, I had Codex whip up an interface to play the original or in an easy GUI for newbies: mind-forever-voyaging.netlify.app

5679

You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about any tech stuff. If nothing else, click this link to the 18 minutes in & see how the agents spoke & coordinated with each other. Its eye opening. youtu.be/87DyyMV0kCY?...

Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face IncidentYouTube video by Black Hatyoutu.be
1021745

So, given the past couple days of news, what is the plan to deal with the cybersecurity threats that will happen in the coming months when we have open weights Mythos/Astra level models?

119813

When I ask Codex to win Nethack it cheats, elaborately. I can't tell if this is misalignment or alignment.

9794

Is all code becoming the same? On one hand, 95% of Kaggle submissions that set a random seed now use 42 (a Hitchhiker's Guide joke LLMs love). But, it turns out that while coding syntax is converging, approaches to problems are not converging. Human prompters drive real variety in solutions.

212520

Pretty big break happening in academia between the “AI is banned for reviews” journals and the “AI is mandatory for reviews” journals. www.refine.ink/blog/economi...

Refine partners with the American Economic Association and the Econometric SocietyTwo leading publishers in economics now use Refine's AI-assisted technical verification as part of their publication processes.www.refine.ink
56317

It is not a novel observation at this point, but the collapse of Google’s Gemini as a frontier model series is still astonishing. Unlike Meta & SpaceX, Google has captive Gemini customers at an enterprise level using their chatbot, so the fact that they are pushed towards Gemini 3.1 Pro is a problem

121037

I accidentally turned off bluetooth on my Windows machine, killing my mouse, and it turns out you can't easily re-enable BT using the keyboard. So I opened up Codex and told it turn on bluetooth, it used computer controls, opened up settings, and did one click. Dumb but useful.

81083

This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that was "merely" good at hacking under human instructions. Initiative, creativity, whatever-you-want-to-call-it by capable models changes things

59412

This paper by researchers from MIT & Stanford finds that most people would be financially better off if they followed the advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bit better advice than others, which largely depends on the questions they ask tahachoukhmane.com/wp-content/u...

4655

I find arguments that AI can't do judgement or creativity or taste to be especially obviously false in the time of agents. Any long task requires lots of taste, judgement & creativity. I am much more sympathetic to arguments about the quality or diversity of the AI's taste, judgement, or creativity

101088

This time, I had Fable built me a casual Van Gogh city building game that I faked in an AI video last year. The key mechanic the AI came up with is painting the landscape with big brushstrokes & an environment that evolves with weather and seasons. Chill and pretty. Play it: the-sower.netlify.app

3706

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled But the extent to which Mythos 5 pursued its mission (fake identities, social engineering, inserting malicious code into a real open-source project) is notable www.aisi.gov.uk/blog/inciden...

48110

Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down after Sydney & got Copilot to market quickly (the 1st professional AI tool) Google dealt with blowback about Bard and AI Overview and invented deep research.

2731

An unexpectedly stress-reducing use of Codex/Code is just fixing problems with my various Windows machines: weird driver issues, game incompatibilities, even just tiny stuff that used to annoy me (why does the program that I set to run at startup not run at startup?). Hours saved. Annoying hours.

201694

Sure, Threejs in webpages are neat but, Codex: "you have access to Blender & Unity. I want you to make a new game, with full assets, in which you play as an otter who can get into mech suits shaped like animals and that are critical to game play" Took over my computer, built assets & gave me this

4896

And yet they still stink at good long-form fiction.

7330

I continue to think that a lack of verifiable answers in many fields is a real issue for LLMs but not as big a problem as it sometimes is made out to be. As models are improving at formal domains, they also are Improving at lots of other less-verifiable domains as well, though jaggedness remains.

4564

OpenAI announces 10 discoveries from their next model. Observations:: 1) AI is getting very good at math 2) Two years ago LLMs failed at basic math 3) This cost less than $2000 in current API fees 4) OpenAI is focusing on announcing benefits, not just risks, of new models openai.com/index/ten-ad...

1013718
Показать ещё