Call to ban AI commits

The explanation for why I think that way are the two paragraphs of preamble. To maybe illustrate this more explicitly, lets take apply the example of curl I provided. They got preferential contact of multiple AI security research groups/teams.

If nix bans AI even if only for commits (which is not necessarily the extend of those arguing for the ban, since the stated ethical concerns apply to any AI use period), Anthropic will under no circumstances provide access to their (newest) model for security audits. Similarly most of those other groups, will see the policy and decide that they don’t want to contribute, since their expectation coming from such a policy is that AI is seen negatively by the project and will look at other projects.

Even if the policy only bans AI commits, that signals to everyone that we are contra-AI, which is something that has been explicitly stated by someone in this thread as one of the goals of such policy. While not strictly banning AI security research, I warn that it could still have similar effect.

Thanks for the reply. Hm. I think people do all sorts of things differently when they’re doing security research, and I’d expect the use of AI to be one of those things, but I could be wrong.

At least in my opinion, if someone were to use a completely local LLM which has had its training data published and has been retrained independently, then id have much less issues with its use. However, do consider that an LLM trained on GPLv3 licensed code is (“should”, whether the lobbied courts agree is a different discussion) be a derived work of said GPLv3 code, as such it itself and any output it produces (which is (should be) itself a derived work) is (should) be licensed under the GPLv3. Which makes even local LLMs very impractical as you’d have to carry around attribution if trained purely on permissive licensed works, or relicense to GPLv3 to be able to utilize GPLv3 training data. As such I don’t view local LLMs (with training data) as practical and local LLMs without training data are still likely to be infringing copyright.

5 Likes

as i believe that everything that happen in a local Git repository is irrelevant and does not need to be moderated, given it is 1. Never shared online and 2. How someone prefer to organise their work is their personal choice

I don’t care what happens on your computer, its your computer, you’re free to do whatever you want with it. But if you have an LLM write code on your computer and then you publish that code online, suddenly it becomes my problem too. Almost all LLMs are proprietary and under honest and proper copyright law should be infringing so much copyright it’s not even funny. As such their use affects the output’s copyright, along with other moral and practical concerns regarding the output. This taints the output with so much baggage that me and others view as immoral, that it leads us to wanting to reject any and all LLM involvement.

9 Likes

There is now another poll. LLMs in nixpkgs

4 Likes

Which makes even local LLMs very impractical as you’d have to carry around attribution if trained purely on permissive licensed works, or relicense to GPLv3 to be able to utilize GPLv3 training data. As such I don’t view local LLMs (with training data) as practical and local LLMs without training data are still likely to be infringing copyright.

So a model like Olmo2, which does not seem to use copyleft data in its training set and discloses the model weights, the methodology and the training data would be fine? I say seem, because I didn’t go through every training data subset included to check whether they fit the license criteria, but a few samples which to me looked like they made a serious effort to 1) prune copyleft data and 2) removed code after removal requests.

Copyleft data seems to be completely incompatible with AI training in general, if your assumption of it infecting the model holds. Sadly I have not found a definitive statement on this, even the FSF is very vague on use of such a license in the training data. It would also make such an AI incompatible with nix repos.

Finally you claim that

At least in my opinion, if someone were to use a completely local LLM which has had its training data published and has been retrained independently, then id have much less issues with its use.

Is that a figure of speech or is there a genuine problem left that hasn’t already been discussed?

1 Like

So a model like Olmo2, which does not seem to use copyleft data in its training set and discloses the model weights, the methodology and the training data would be fine?

If what they state is true, then such a model would in my mind be okay. I probably won’t use it myself due to my general built up hate towards LLMs.

Honestly I’m not even sure what I would use an LLM for. But thank you nonetheles.

Is that a figure of speech or is there a genuine problem left that hasn’t already been discussed?

I still have reservations about the energy use, infrastructure cost, social costs among others, but at least, for me, one of the more personal gripes is resolved with models such as that one.

4 Likes

This thread is looooong and I wanna have an idea on what is the more popular position in this community

  • What’s your opinion?
  • All AI contributions must be banned.
  • Allow AI contributions but limit the scope and add more code-review scrutiny.
  • Allow AI contributions and don’t change much about how maintaining projects go.
0 voters

@ammaratef45 I don’t know if recontinuing this thread is the best idea :smiling_face_with_tear: but also, doesn’t the poll in this thread’s very first post already answer your question?

10 Likes

urgh sawwy, I didn’t see the poll in the post somehow, closed mine

1 Like

Except for dynamic derivation or runtime Nix, humans must verify Nix code and commit it; if we leave this to AI for nixpkgs, I believe the community is finished. nix is ​​not lisp.

2 Likes

Uhm, Nixpkgs uses a lot of deterministic generators with quite diverse cases where the scope is too large for meaningful manual review.

5 Likes

nix is ​​not lisp.

Even if nix was a lisp, i personally dont want LLMs involved.

5 Likes