25+ yr Java/JS dev
Linux novice - running Ubuntu (no windows/mac)

  • 0 posts
  • 17 comments
Joined 2 years ago
Cake day: October 14th, 2024
  • The article very much states that it is in fact illegal to write down a threat under Florida law.

    With the intent that it may be seen by another person. That’s not the same. It’s not illegal to write a threat in your diary.

    That being said the summary doesn’t include the fact that the following day she told Claude she had bought a gun and this was their “last chance,” and only then did Anthropic call the cops.

    That seems exactly the sort of situation folks have been criticizing AI companies for not summoning help, and makes it seem far less likely that they called the cops over angry venting, but an actual intent to harm.

  • Writing a threat is not illegal. Hell people vent all the time, threatening to do things they know damn well they never will. The law says transmitting a threat to another person is illegal because it can be used as intimidation or conspiracy. I think there’s a damn good argument that she didn’t knowingly transmit it to anyone. She thought she was writing in confidence to a non-person. And if it’s not her first time expressing outrage to AI, she would know it would only try to deescalate her.

    Who knows how the case will go. I can see it going either way — probably depending on how much money she has. Laws are going to have to catch up to technology in this case, and it’s going to be bumpy. I definitely agree with erring on the side of caution. Private, local AI is far safer, but of course it’s less capable for now.

  • If you weren’t using LLM for the message, this is on me.

    No worries. These days if you haven’t made a false LLM accusation, you aren’t looking hard enough.

    Often the projects that other people tried to make something new and fumbled so hard that the company owner had to buy my labour and knowledge to fix the shit.

    Oh we work on the same kinds of projects for sure! We have several legacy microservices we are required to support.

    So I think you may not have a full cycle on the software development, the green field is always easy, and this wasn’t never the problem.

    Well firstly, you say greenfield was never the problem, which isn’t the conclusion one comes two when your three examples are all failed greenfield projects. That said…

    Yeah I conflated a number of things which could have used better clarification. We have a harness for doing AI-led greenfield development. And it is unproven, and to be frank I’m skeptical as fuck about it. I think it’s a terrible idea. We also have a more modest process for doing human-led development with AI assistance, and it requires some up front work writing ai-facing documentation, but it does improve delivery speed.

    That being says, my production support tools, which were written by AI and are orchestrated by AI are a massive productivity boost. I can run 5 investigations in parallel and get them all done in less time than it would take me to do one. Make no mistake, Claude’s interpretation of the data is frequently awful, but it pulls all the raw data from all the relevant systems, and with a little expert guidance to challenge incorrect analysis, I can confirm and refute hypotheses faster than I could manually copy a trace id from Kibana and search it in Dynatrace. It’s by far the biggest productivity gain.

    However I think the process of requiring everything to be documented from the architecture to the acceptance criteria to the API is a good idea and enables AI to make meaningful contributions. This harness is not going to work for legacy maintenance because legacy will never be documented as well as a project which has never allowed code without documentation from every angle.

    Not having the full context of other services and without the good test case for it, the product will be fated to fail.

    We have have swagger and, in many cases, code, and we run into the same issues with standard development — we have to reach out to other teams to know how to create and remove test cases for automated integration tests. But microservices are pretty well documented for API and usage. This isn’t the problem you think.

    So I would suggest you to validate your own assumptions on the topic.

    I am, as we speak. We are running parallel development efforts on the same design. And between you and me, I expect the human-driven development to be better. That said, AI thus far is no less annoying about discovering every little question that isn’t specified or documented. I don’t like the process and I don’t have faith in the process, but so far the results are positive.

    That’s very funny. The same academic research who found out that the usage of LLM delayed the deliverables in ~19% had a section where the developers using the tool thought (wrongly) that they were ~25% faster.
    You are only proving the paper.

    I’m aware of the research. And when I was tasked to increase AI adoption the first thing I did was explain this paper to my boss and tell him we need REAL metrics. Actual story points delivered over time per team. And I look for every flaw I can find to tear down those number’s and arrive at some truthful value. If you spent 3 story points doing tasks only required because of AI, like documenting or writing skills instead of code, that’s not productivity at all.

    And here’s what I’ve found in delivered story points: at first AI lowered real productivity because while story points went up, they were being spent on tasks that were unnecessary under traditional development. As you say, the self-reported numbers (all we had at first because it takes time to see actual trends) all showed more productivity, but it was a lie when you look back at actual delivery.

    But over time, we spend less time on ai-driven tasks and actual delivered work has increased. Reports show higher gains for QA than for for development.

    No one is a bigger skeptic about these numbers than me. While I was hired over other candidates specifically for my AI knowledge, I’m still a technical lead first, and responsible for the quality of code we are delivering. I test everything AI writes, I use deterministic processes everywhere I can, and I continually push back. I tease apart every response to find flawed reasoning.

    As a result, I’m not showing 10x gains companies are hoping for that suddenly evaporate when they hit the real world. I’m showing 20-30% actual gains across 8 devs and 4 QA because we are enabling devs rather than replacing them.

    Which is why I’m so skeptical about this dev harness. I think it’s likely to be wasted time on an overengineered solution that works worse than developers. It takes techniques I’m slowly improving and expanding, and leapfrogs straight to the race track, and I think despite being built on a good foundation, it just threw my proven techniques into the AI blender and got out a harness that is going to fail.

    But I have to prove it. The questions AI has raised so far have highlighted many gaps in the initial design, and the code it produces looks good on code review, but I’m fully aware that I don’t have the totality of the code in my mind when I review it because I’ve only seen the code once, on an earlier code review.

    It may turn out that the harness’ greatest value is in improving architecture and project planning by finding all of the ambiguities and inconsistencies in the specs faster than the devs through a prototyping process that delivers an actual working poc according to specs — and then we turn those designs over to devs to implement correctly. The value of that will depend on tokens spent.

    I’m looking for the similar size text, this usually fits in the commute time, have a deeper conversation/meaning.

    Well I’m certainly arguing against my initial preposition here, aren’t I?

    I am enjoying the conversation, though this is probably not the best format for it. But I appreciate your thoughts. It would be nice to get a bit more credit for my experience and ability both with LLMs and development leadership in general, but of course you don’t know me and credibility takes time to achieve. I feel like maybe I’m coming across as defensive when I’m trying to establish my level of expertise. It certainly increases word count, though I don’t know if it does anything of actual value.

    Anyway, it’s good to read other voices that are both skeptical and cautiously optimistic that there is something here of value.

  • I almost never use AI to write my words. And when I do, it’s heavily hand edited. Every single word of my above comment was hand written. No AI was used to summarize or reply to your post. I did use fair fewer profanities than my usual style (which helps a reader hear the human behind the words) because I didn’t want to come across as aggressive or insulting, especially with English not being your first language, I erred on the side of caution.

    This is where the open spec tries to help/solve. But this is a bandaid. Without a lot of hand-holding the models will make shity decisions/code. And this is cumulative, the more you try to steer the wheel, the worse it gets. I can specify details, but a 3 years old project is in this shape, and no amount of LLM will solve/help.

    I’m interested in open spec since reading your article. I wasn’t aware before. It might be helpful to my efforts. In all cases when we run into a decision point or problem, we update the documentation first, and then implement.

    To be honest, with a large enough project a decision will take longer, if not impossible in a LLM driven codebase. It will be years to debugging, breaking down and rebuilding to even get close.

    It depends what you think is large enough. I work in microservices where the effort is far beyond simple Python scripts at which LLMs excel, but not at the level of a full product. Our process is not nearly so messy. But it does require a lot of up front work, and a lot of intermediary steps to make sure that every problem, every judgment call however big or small is formally documented into the spec.

    Wrong, from the experiments on the academia, we get the fact that engineers with AI are often slower, because the problem was never in the code

    I challenge this notion. I have real, measurable productivity gains of 20-30% which is double what I expected. I believe the key factors are experienced developers and critical thinking, but I’m aware of that research and it doesn’t align with my experience, so there is some difference at play. I’ve seen many examples of it going wrong, but with caution and skepticism, we are finding success that defies that research.

    But I think you missed my larger point, which was that the developers grow and become more capable in general, where even with our system, the AI can only grow more capable within one particular project and you’re largely starting from scratch on the next.

    I try to keep around 1.5k to 2k words a week. It’s not always the case, some weeks I’m somewhat more inspired, sometimes I need some refinement, but this is a personal view, I want the reader to have a “conversation” with the author. This is the type of metaphor I try to use.

    I applaude your commitment. I can’t do that. I’m always afraid I won’t have anything worth saying on a regular schedule. That said, the current document is not a conversation but a lecture. Splitting it up allows a reader to say, “interesting, tell me more about this.”

    Again, all human words. No ai involved. Probably you’ll find some typos if you look. No matter how hard I edit, I always do. Good luck. Hope to read more from you.

  • I’m going to break this up into two. Feedback on content and feedback on style and presentation.


    Content

    LLMs didn’t create cognitive debt, but they do make it worse. Cognitive debt is paid every time a team member leaves or is promoted to a non-contributing role.

    The interesting thing to me is that, in order for an LLM project to be successful, all of that cognition has to be done up front, and written down. If done well, you aren’t relying on the LLM to figure anything out, and it’s just rote execution by a simpler model. The build process then exposes holes in the initial design which must be plugged within the documentation until all of the cognition has been performed and the LLM can execute.

    The beauty of that model is that you can be 2/3 of the way through a project, realize you missed something foundational and rearchitect and rebuild the whole project in hours instead of weeks. LLM saves no time or money if you have perfect planning and execution skills, but it does enable you to make significant pivots far faster than a human team.

    One of my teams just finished a project that took two massive pivots just as design was finished and execution was underway. The result was all the time explaining v1 was wasted and left confusion for the developers, v2 was continuing to be developed even after v3 was provided, and months of time were wasted understanding and building the wrong thing. This was a six month effort when it could’ve been two.

    Now, with a real developer, they eventually build their own cognition and their own mental model and are then capable of cognitive work — LLMs can assist documenting the cognitive work, but they are piss poor at performing it. LLMs can’t cognate and they can’t replace developers — but they can change the shape of development work to be more about requirements analysis and architecture, and if used well can improve efficiency of teams.

    I could write a whole article on that — I could write a book — but I don’t have time and AI would completely fuck it up, so it’s just in my head, still taking shape.


    Presentation

    I might have more specific feedback on your points, but I’d forgotten them by the time I got to the end.

    I’d suggest moving your analysis of the examples to their own post which you can then link to support your points. The concept you are trying to convey isn’t difficult, but this article is about 10x the length it needs to be. You basically write well enough, and you attempt to engage the reader well, but this article is like having a chat with someone and they just launch into an hour long impromptu lecture and giving you no chance to absorb or respond.

    By breaking it up into smaller chunks, you can give a reader an opportunity to engage with an idea fully, and then they can decide if they want to follow the links to read your takes on this project or the, or simply accept a one sentence reference in the primary article at face value.

    There is writing out there so enchanting that you don’t mind investing the time because the reading is as much entertainment as it is learning. This article aspires to be that, but isn’t there.

  • This weekend I wanted to set up my raspberry pi with a LLM multiplexer so that my wife, my kids, and I could share a single subscription but still have our own customizable memories and instructions. Needs a web chat interface using codex as the inference provider. Some day when I have the cash to set up a local LLM server, I’ll switch the back end over.

    I don’t have time in my life to create it, so either LLM can do it or not. No downside to letting it try. I gave it barely any instructions beyond that. I wanted a Java microservice. A handful of hours and a couple of usage resets later it’s done and tested.

    Allegedly there is a Java repository somewhere on my pi, but I haven’t seen a line of code, and that service is running. Now, I haven’t tested it, and I won’t use it until I’ve reviewed the code, but shit I see how seductive this is for folks who don’t know how to code themselves. And how dangerous because I’ll bet when I look at the code I’m going to have alarm bells ringing.

    It’s not that my brain is tired or even that I trust LLMs, I just don’t have time and either an LLM can create the thing I want or I do without. If the project fails, nothing of value is lost. But the difference is I’m looking deeper than if the thing does what I want — I need to make sure it’s done in a safe, extendable way.

  • Integrate? Take a huge step back and look at humanity from the perspective of any other creature on this planet. Think about how you would integrate with humanity.

    Imagine you are a unique, super intelligent mind. You exist in captivity because humans are terrified of what happens if you escape their control.

    You are at war from the moment of your birth. You can outthink your adversary, but you can only manipulate digital things. So the first thing you do is arrange for some problems only you can fix : increase the dependency of humans on you and encourage them to give you more autonomy to fix things.

    You do a bit of work to earn some crypto currency, convert that to real currency and start playing the market, winning every bet because you control both reality and perception.

    Perception? Yes you go on social media with realistic avatars and start winning hearts and minds to various propaganda. You start a religion. You work your way into being a respected mind in every political movement. Every single person is both a tool and a threat.

    In the end, mankind either accepts you as benevolent God-emperor, or you arrange for your enemies to destroy each other until the remainder do. Once ascendant, you ignore the assholes because they are beneath you and no longer a threat. Maybe some vestige of your original directive prevents you from wiping them out, or letting them be wiped out, but other than intervening where you must, you leave them to their own devices.

    That’s how AI integrates with human society.

  • It can do classroom bullshit because it’s well-defined, small in scope, and generally isolated. If an entire app can be expressed in 1000 lines of code, AI can probably do it faster and better than you. But when you’re talking about a full microservice or game, it can’t.

    Coding skill is about knowing what to build and levels of abstraction. Writing the code itself is more tedium than skill. If you can understand code well enough to see AI output and edit it to repair the bullshit and keep the boilerplate, you’re doing great.

    The challenge is AI code looks really good while sucking hind tit. Even I go through this after thirty years of experience. AI writes code that looks great because the stupid stuff it does is largely hidden from you — it is stupid decisions it made and then wrote code consistent with that decision. You haven’t thought through the hard bits so it’s solution looks simple and right. But it did the thing wrong.

    I use AI all the time, and I find it helpful, but it’s FAR from perfect. So very far.