As we continue developing our software, we accumulate a growing amount of technical debt just to keep the system running. But I believe we are on the brink of an even larger issue. Cognitive debt.

Hope you enjoy this reading, all feedback is welcome.

  • fruitycoder@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    2
    ·
    16 hours ago

    I hate to say this but can’t LLMs solve the cognitive debt problem better then people? Wouldn’t having a stochastic machine that can have its state frozen in time, retrived at convience, and fed the same inputs be a potentially more transparent machine, thinking or otherwise?

    I agree with you on stating thr problem though, and the term seems concise enough to me. The gap seems to be not in ability to do this but that creating graspable and reasonable audits for LLM usage is like many risk mitigation systems, an after thought in the industry.

    • wholookshere@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      5
      ·
      12 hours ago

      and fed the same inputs be a potentially more transparent machine, thinking or otherwise?

      This is where you go off the rails. LLMs are probalistic, not deterministic. It will probably give the same output, but thats still not actually true. This is where hallucinations come from.

      • fruitycoder@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        1
        ·
        10 hours ago

        Tbf the exact input and node activation is all you’d need for forensics, the only reason to refeed would old input plus new input mix, which should be new output.

            • wholookshere@lemmy.blahaj.zone
              link
              fedilink
              English
              arrow-up
              1
              ·
              9 hours ago

              For training maybe. But execution as far as i know are just a bunch of probabilities chucked into matrices. Back in my physcos days I could u derstand the math, but not today.

              Also what needs citing, that nodes translate to debugable, reproducible outputs.

              Because again, it’s all probablilties under the hood.

              • fruitycoder@sh.itjust.works
                link
                fedilink
                English
                arrow-up
                1
                ·
                8 hours ago

                I mean probabilities in matrices are nodes in a neural net, right?

                The only thing that makes it non-deterministic is the tempature value which is known after the fact from my understanding, so that should be able to deterministic.

                • wholookshere@lemmy.blahaj.zone
                  link
                  fedilink
                  English
                  arrow-up
                  1
                  ·
                  edit-2
                  8 hours ago

                  No… No they’re not. Either that or nerual nets are even dumber than I thought. My understanding was emulating nodes of information like brains do. Thats not probabilities. But willing to be wrong if you can bring sources. The connection between points in probability doesn’t make sense. Thats a nonsense statement.

                  Thats also no how simulated annealing works. The temperature value is roughly the probability it will pick a different, less optimal step, in order to try and find better alternative paths. The temperature value is roughly the probability of trying something ‘random’. Not deterministic at all. The temperature value is lowered with time and progression. Its known the entire time, but its still a weighted coin flip that determines which path to take.

                  https://en.wikipedia.org/wiki/Simulated_annealing

  • Prox@lemmy.world
    link
    fedilink
    English
    arrow-up
    7
    ·
    2 days ago

    This is the cognitive debt of the past. We already had it. We just didn’t have the name for it.

    We always called it “domain knowledge” or “tribal knowledge”.

    • Shin@piefed.socialOP
      link
      fedilink
      English
      arrow-up
      3
      arrow-down
      1
      ·
      22 hours ago

      But in the days of old, the replacement of human, or generation of domain/tribal knowledge were in the range of human capability to learn the new parts of a system.

      With the LLM of today this scale is far flipped on the wrong side, and no product fully LLM will be manageable in a few years.

    • corsicanguppy@lemmy.ca
      link
      fedilink
      English
      arrow-up
      6
      arrow-down
      2
      ·
      2 days ago

      “oral history”.

      “The ropes”

      It’s got many names. Why did this guy not know any of them? Why was “I don’t know the name” suddenly “there is no name”?

  • AnarchistArtificer@lemmy.world
    link
    fedilink
    English
    arrow-up
    12
    ·
    2 days ago

    I enjoyed reading this. I don’t have anything more constructive to say (I’m only leaving this comment because I realised that you are the person who actually wrote the linked post, and I’ve found that leaving an appreciative comment is like a super-upvote).

    I like this piece because it’s another voice saying effectively the same thing as many others I respect have argued: that pressure to ship leads to poor quality, difficult to maintain code (even before AI), and that AI is just making the problem worse. It’s cool to see the conversation evolve. For example, after I first saw the term “cognitive debt” used, I saw it crop up in more and more places, as other writers seemed to conclude “this problem is similar to technical debt, but it’s distinct enough that using a different but related term is a good idea”. And even if there are multiple people arguing effectively the same thing, each person tends to tackle the problem slightly differently

    • Shin@piefed.socialOP
      link
      fedilink
      English
      arrow-up
      3
      ·
      2 days ago

      Hey there! Thanks for the comment, and reading it.

      I know that sometimes hearing the same voice speaking on the same topic can be problematic. But on the other side of the same coin, have fewer voices pointing to the problem is also a problem. For me, writing the blog is a way to organize my own mind and be able to create a reasoning on some topic.

      And sometimes, repeating the echoed words that the industry is repeating.

      Either way, thanks again for reading, and even more for the comment.

    • Shin@piefed.socialOP
      link
      fedilink
      English
      arrow-up
      2
      ·
      22 hours ago

      I’ll be honest, I had to look on the dictionary to understand this subtle meaning.
      Thanks, I learned a new thing today! :D

  • MagicShel@lemmy.zip
    link
    fedilink
    English
    arrow-up
    2
    arrow-down
    4
    ·
    2 days ago

    I’m going to break this up into two. Feedback on content and feedback on style and presentation.


    Content

    LLMs didn’t create cognitive debt, but they do make it worse. Cognitive debt is paid every time a team member leaves or is promoted to a non-contributing role.

    The interesting thing to me is that, in order for an LLM project to be successful, all of that cognition has to be done up front, and written down. If done well, you aren’t relying on the LLM to figure anything out, and it’s just rote execution by a simpler model. The build process then exposes holes in the initial design which must be plugged within the documentation until all of the cognition has been performed and the LLM can execute.

    The beauty of that model is that you can be 2/3 of the way through a project, realize you missed something foundational and rearchitect and rebuild the whole project in hours instead of weeks. LLM saves no time or money if you have perfect planning and execution skills, but it does enable you to make significant pivots far faster than a human team.

    One of my teams just finished a project that took two massive pivots just as design was finished and execution was underway. The result was all the time explaining v1 was wasted and left confusion for the developers, v2 was continuing to be developed even after v3 was provided, and months of time were wasted understanding and building the wrong thing. This was a six month effort when it could’ve been two.

    Now, with a real developer, they eventually build their own cognition and their own mental model and are then capable of cognitive work — LLMs can assist documenting the cognitive work, but they are piss poor at performing it. LLMs can’t cognate and they can’t replace developers — but they can change the shape of development work to be more about requirements analysis and architecture, and if used well can improve efficiency of teams.

    I could write a whole article on that — I could write a book — but I don’t have time and AI would completely fuck it up, so it’s just in my head, still taking shape.


    Presentation

    I might have more specific feedback on your points, but I’d forgotten them by the time I got to the end.

    I’d suggest moving your analysis of the examples to their own post which you can then link to support your points. The concept you are trying to convey isn’t difficult, but this article is about 10x the length it needs to be. You basically write well enough, and you attempt to engage the reader well, but this article is like having a chat with someone and they just launch into an hour long impromptu lecture and giving you no chance to absorb or respond.

    By breaking it up into smaller chunks, you can give a reader an opportunity to engage with an idea fully, and then they can decide if they want to follow the links to read your takes on this project or the, or simply accept a one sentence reference in the primary article at face value.

    There is writing out there so enchanting that you don’t mind investing the time because the reading is as much entertainment as it is learning. This article aspires to be that, but isn’t there.

    • Shin@piefed.socialOP
      link
      fedilink
      English
      arrow-up
      8
      ·
      2 days ago

      Wooow, this message is as long as my post, love it!

      The interesting thing to me is that, in order for an LLM project to be successful, […]

      This is where the open spec tries to help/solve. But this is a bandaid. Without a lot of hand-holding the models will make shity decisions/code. And this is cumulative, the more you try to steer the wheel, the worse it gets. I can specify details, but a 3 years old project is in this shape, and no amount of LLM will solve/help.

      The beauty of that model is that you can be 2/3 of the way through a project, […]

      I understand this point, but it’s flaky at best. To be honest, with a large enough project a decision will take longer, if not impossible in a LLM driven codebase. It will be years to debugging, breaking down and rebuilding to even get close. Or spend the paycheck of 5 engineers annually for a single migration.

      And this considers the the current LLM are heavily subsided.

      Now, with a real developer, they eventually build their own cognition and their own mental […]

      Wrong, from the experiments on the academia, we get the fact that engineers with AI are often slower, because the problem was never in the code, code is the tool, it was the acquiring the right data, model and get the knowledge of the product to make the changes needed.

      Presentatio […]

      I try to keep around 1.5k to 2k words a week. It’s not always the case, some weeks I’m somewhat more inspired, sometimes I need some refinement, but this is a personal view, I want the reader to have a “conversation” with the author. This is the type of metaphor I try to use.

      Either way, thanks for reading.

      Ps.: Do you always use LLMs to write your messages? I would love to read your own words. Don’t be the meat proxy for the LLM.

      • MagicShel@lemmy.zip
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        1 day ago

        I almost never use AI to write my words. And when I do, it’s heavily hand edited. Every single word of my above comment was hand written. No AI was used to summarize or reply to your post. I did use fair fewer profanities than my usual style (which helps a reader hear the human behind the words) because I didn’t want to come across as aggressive or insulting, especially with English not being your first language, I erred on the side of caution.

        This is where the open spec tries to help/solve. But this is a bandaid. Without a lot of hand-holding the models will make shity decisions/code. And this is cumulative, the more you try to steer the wheel, the worse it gets. I can specify details, but a 3 years old project is in this shape, and no amount of LLM will solve/help.

        I’m interested in open spec since reading your article. I wasn’t aware before. It might be helpful to my efforts. In all cases when we run into a decision point or problem, we update the documentation first, and then implement.

        To be honest, with a large enough project a decision will take longer, if not impossible in a LLM driven codebase. It will be years to debugging, breaking down and rebuilding to even get close.

        It depends what you think is large enough. I work in microservices where the effort is far beyond simple Python scripts at which LLMs excel, but not at the level of a full product. Our process is not nearly so messy. But it does require a lot of up front work, and a lot of intermediary steps to make sure that every problem, every judgment call however big or small is formally documented into the spec.

        Wrong, from the experiments on the academia, we get the fact that engineers with AI are often slower, because the problem was never in the code

        I challenge this notion. I have real, measurable productivity gains of 20-30% which is double what I expected. I believe the key factors are experienced developers and critical thinking, but I’m aware of that research and it doesn’t align with my experience, so there is some difference at play. I’ve seen many examples of it going wrong, but with caution and skepticism, we are finding success that defies that research.

        But I think you missed my larger point, which was that the developers grow and become more capable in general, where even with our system, the AI can only grow more capable within one particular project and you’re largely starting from scratch on the next.

        I try to keep around 1.5k to 2k words a week. It’s not always the case, some weeks I’m somewhat more inspired, sometimes I need some refinement, but this is a personal view, I want the reader to have a “conversation” with the author. This is the type of metaphor I try to use.

        I applaude your commitment. I can’t do that. I’m always afraid I won’t have anything worth saying on a regular schedule. That said, the current document is not a conversation but a lecture. Splitting it up allows a reader to say, “interesting, tell me more about this.”

        Again, all human words. No ai involved. Probably you’ll find some typos if you look. No matter how hard I edit, I always do. Good luck. Hope to read more from you.

        • Shin@piefed.socialOP
          link
          fedilink
          English
          arrow-up
          2
          ·
          22 hours ago

          First, I’m sorry. I attribute LLM use when I had only suspicions about it. And this is on me. If you weren’t using LLM for the message, this is on me.


          This is where the open spec tries to help/solve […]

          I’m interested in open spec since reading your article. I wasn’t aware before. […]

          This is a derivate methodology from the SpecKit (which is a piece of crap, stay away from it). It helps to certain points, but this is still bandaid in it. Not a real solution.

          It depends what you think is large enough. I work in microservices where the effort […]

          Looks like we work in very different kind of projects. I’m used to jump into garbage projects to fix it. Often the projects that other people tried to make something new and fumbled so hard that the company owner had to buy my labour and knowledge to fix the shit.

          Yet, microservice is the point were LLMs fails the most. Not having the full context of other services and without the good test case for it, the product will be fated to fail. So I think you may not have a full cycle on the software development, the green field is always easy, and this wasn’t never the problem.

          Look on the history, the first months of a product is always the most productive (way before any LLM would ever exist), the real issues appears years in the development, when the cumulative decisions start to group and becames a problem, where every new change breaks other. The technical debt that I spoke in the article.

          This isn’t on the first week of the project, this is years in it.

          In short, microservice is already a bad design for most of products, very few products require to be microservices, and combined with the LLM lack of view on the product as a whole makes this even worse. So I would suggest you to validate your own assumptions on the topic.

          There is a good chance that you are either not fully validating, or not seeing the product as a whole.

          Wrong, from the experiments on the academia, […]

          I challenge this notion. I have real, measurable productivity gains of 20-30% which […]

          But I think you missed my larger point, which was that the developers […]

          That’s very funny. The same academic research who found out that the usage of LLM delayed the deliverables in ~19% had a section where the developers using the tool thought (wrongly) that they were ~25% faster.

          You are only proving the paper.


          Regarding the idea, I’ll think about it (like I always do), but making shorter text will not make my style. When I’m reading blogs I’m looking for the similar size text, this usually fits in the commute time, have a deeper conversation/meaning.

          Again, thanks for reading.

          • MagicShel@lemmy.zip
            link
            fedilink
            English
            arrow-up
            1
            ·
            18 hours ago

            If you weren’t using LLM for the message, this is on me.

            No worries. These days if you haven’t made a false LLM accusation, you aren’t looking hard enough.

            Often the projects that other people tried to make something new and fumbled so hard that the company owner had to buy my labour and knowledge to fix the shit.

            Oh we work on the same kinds of projects for sure! We have several legacy microservices we are required to support.

            So I think you may not have a full cycle on the software development, the green field is always easy, and this wasn’t never the problem.

            Well firstly, you say greenfield was never the problem, which isn’t the conclusion one comes two when your three examples are all failed greenfield projects. That said…

            Yeah I conflated a number of things which could have used better clarification. We have a harness for doing AI-led greenfield development. And it is unproven, and to be frank I’m skeptical as fuck about it. I think it’s a terrible idea. We also have a more modest process for doing human-led development with AI assistance, and it requires some up front work writing ai-facing documentation, but it does improve delivery speed.

            That being says, my production support tools, which were written by AI and are orchestrated by AI are a massive productivity boost. I can run 5 investigations in parallel and get them all done in less time than it would take me to do one. Make no mistake, Claude’s interpretation of the data is frequently awful, but it pulls all the raw data from all the relevant systems, and with a little expert guidance to challenge incorrect analysis, I can confirm and refute hypotheses faster than I could manually copy a trace id from Kibana and search it in Dynatrace. It’s by far the biggest productivity gain.

            However I think the process of requiring everything to be documented from the architecture to the acceptance criteria to the API is a good idea and enables AI to make meaningful contributions. This harness is not going to work for legacy maintenance because legacy will never be documented as well as a project which has never allowed code without documentation from every angle.

            Not having the full context of other services and without the good test case for it, the product will be fated to fail.

            We have have swagger and, in many cases, code, and we run into the same issues with standard development — we have to reach out to other teams to know how to create and remove test cases for automated integration tests. But microservices are pretty well documented for API and usage. This isn’t the problem you think.

            So I would suggest you to validate your own assumptions on the topic.

            I am, as we speak. We are running parallel development efforts on the same design. And between you and me, I expect the human-driven development to be better. That said, AI thus far is no less annoying about discovering every little question that isn’t specified or documented. I don’t like the process and I don’t have faith in the process, but so far the results are positive.

            That’s very funny. The same academic research who found out that the usage of LLM delayed the deliverables in ~19% had a section where the developers using the tool thought (wrongly) that they were ~25% faster.
            You are only proving the paper.

            I’m aware of the research. And when I was tasked to increase AI adoption the first thing I did was explain this paper to my boss and tell him we need REAL metrics. Actual story points delivered over time per team. And I look for every flaw I can find to tear down those number’s and arrive at some truthful value. If you spent 3 story points doing tasks only required because of AI, like documenting or writing skills instead of code, that’s not productivity at all.

            And here’s what I’ve found in delivered story points: at first AI lowered real productivity because while story points went up, they were being spent on tasks that were unnecessary under traditional development. As you say, the self-reported numbers (all we had at first because it takes time to see actual trends) all showed more productivity, but it was a lie when you look back at actual delivery.

            But over time, we spend less time on ai-driven tasks and actual delivered work has increased. Reports show higher gains for QA than for for development.

            No one is a bigger skeptic about these numbers than me. While I was hired over other candidates specifically for my AI knowledge, I’m still a technical lead first, and responsible for the quality of code we are delivering. I test everything AI writes, I use deterministic processes everywhere I can, and I continually push back. I tease apart every response to find flawed reasoning.

            As a result, I’m not showing 10x gains companies are hoping for that suddenly evaporate when they hit the real world. I’m showing 20-30% actual gains across 8 devs and 4 QA because we are enabling devs rather than replacing them.

            Which is why I’m so skeptical about this dev harness. I think it’s likely to be wasted time on an overengineered solution that works worse than developers. It takes techniques I’m slowly improving and expanding, and leapfrogs straight to the race track, and I think despite being built on a good foundation, it just threw my proven techniques into the AI blender and got out a harness that is going to fail.

            But I have to prove it. The questions AI has raised so far have highlighted many gaps in the initial design, and the code it produces looks good on code review, but I’m fully aware that I don’t have the totality of the code in my mind when I review it because I’ve only seen the code once, on an earlier code review.

            It may turn out that the harness’ greatest value is in improving architecture and project planning by finding all of the ambiguities and inconsistencies in the specs faster than the devs through a prototyping process that delivers an actual working poc according to specs — and then we turn those designs over to devs to implement correctly. The value of that will depend on tokens spent.

            I’m looking for the similar size text, this usually fits in the commute time, have a deeper conversation/meaning.

            Well I’m certainly arguing against my initial preposition here, aren’t I?

            I am enjoying the conversation, though this is probably not the best format for it. But I appreciate your thoughts. It would be nice to get a bit more credit for my experience and ability both with LLMs and development leadership in general, but of course you don’t know me and credibility takes time to achieve. I feel like maybe I’m coming across as defensive when I’m trying to establish my level of expertise. It certainly increases word count, though I don’t know if it does anything of actual value.

            Anyway, it’s good to read other voices that are both skeptical and cautiously optimistic that there is something here of value.