Over the course of the year, it’s become increasin...
# general
v
Over the course of the year, it’s become increasingly clear that writing code is one of the things LLMs are most capable of.
Over the course of the year, it’s become increasingly clear that writing code is one of the things LLMs are most capable of. If you think about what they do, this isn’t such a big surprise. The grammar rules of programming languages like Python and JavaScript are massively less complicated than the grammar of Chinese, Spanish or English. It’s still astonishing to me how effective they are though. One of the great weaknesses of LLMs is their tendency to hallucinate—to imagine things that don’t correspond to reality. You would expect this to be a particularly bad problem for code—if an LLM hallucinates a method that doesn’t exist, the code should be useless. Except... you can run generated code to see if it’s correct. And with patterns like ChatGPT Code Interpreter the LLM can execute the code itself, process the error message, then rewrite it and keep trying until it works! So hallucination is a much lesser problem for code generation than for anything else. If only we had the equivalent of Code Interpreter for fact-checking natural language! How should we feel about this as software engineers? On the one hand, this feels like a threat: who needs a programmer if ChatGPT can write code for you? On the other hand, as software engineers we are better placed to take advantage of this than anyone else. We’ve all been given weird coding interns—we can use our deep knowledge to prompt them to solve coding problems more effectively than anyone else can.
g
is it really? which one? I'm still not getting the amazing code that everyone else seems to enjoy
p
I've had great outputs from ChatGPT Plus (GPT-4) and Cursor (using GPT-4 API). Tabnine, Codeium, Llama weren't that great though.
l
@Gwen Shapira if you can share, what are you using? I've heard this more than once by now and I find it surprising because copilot is very transformative tech. I have some sparse notes about "programming in the age of AI" so maybe this conversation convinces me to finally organise them and share them 🙂
there's also something to say about the fear of AI taking over our jobs of course (which I understand well but I think it's technically unfounded)
m
Copilot is nice to write boilerplate for you. Sometimes it does make really great suggestions where it surprises me. Most of the time though, for complex tasks, it takes me so long to tell Copilot what to do that I could have written it myself in the same time.
g
Yeah, I use co-pilot and like it. But anything over few lines of boilerplate, and its hilariously wrong. Especially when it makes up method, types, etc. I was talking more about the videos / tweets I occasionally see with "I don't know how to code, but I build this [ game / sales app / marketing app / productivity tool ] with GPT4. I was trying to use GPT4 for my demos, and never managed to get anything reasonable. I need to spend hours explaining what I need and then also fix it myself. And if I know enough to fix it, I could probably just write it in the first place...
l
This is so interesting to me! I hear that consistently and I’m amazed at how different from my experience is. Let me try a short tl;dr (I will fail at being short). My experience is very different: boilerplate sure, occasional hallucination also sure. But I also get incredibly accurate results in some specific contexts. It may sound small but it’s so transformative that in fact it already changed my workflow completely. Best example scenario: I write some code and then I write the tests. When I write the tests, copilot can often generate perfect tests just from the test case description. Which makes me almost 30% faster? I can’t tell but feels like that much. I also noticed that the quality of the suggestion is both language and context dependent. In Go, if I write comments for what I want I often get production grade code at first try. In Kotlin, that doesn’t work. But when I’d doing test case description, I get better Kotlin results. One very funny scenario I didn’t consider before starting to use copilot: Avro, protobuf definitions.. I can mostly just describe what I need and ignore the syntax :D As for chats, I can’t integrate them in my workflow yet. It’s so much hit and miss for me at the moment. Something worth underlining: I’m much more impatient about the quality of the results in chat gpt. The media feels more “human” so I expect more? Can’t tell. Sorry, I knew I’d fail the length test 🤦‍♂️