Re: [MAINTAINERS SUMMIT] Other LLM-related topics - tags, newcomers, etc
Mauro Carvalho Chehab <[email protected]> Fri, 17 Jul 2026 09:19:38 +0200
| Newsgroups | dev.linux.lists.ksummit |
|---|---|
| Message-ID | <[email protected]> |
On Thu, 16 Jul 2026 22:27:54 -0400 "Theodore Tso" <[email protected]> wrote: > On Fri, Jul 17, 2026 at 02:58:12AM -0500, Mauro Carvalho Chehab wrote: > > > > The main point is: do we really need 671B parameters? Those models > > speak a lot of different languages, have medical databases, and a lot > > of other random knowledge that are useless for kernel development. > > Fair; but I do think we need to fine-tune it with a lot LKML text so > that it has the knowledge that we actually need for kernel development > before I'd trust the smaller models. > > > I've been playing for a while with qwen 3.6 with 24KB context size, > > 36B parameters (3B activated), 4bits kv quantization and it does produce > > some decent results. The main limitation is the context size: it is > > probably not big enough to test big files (*) > > I've been playing with qwen3-next with 256k context context size, 80B > parameters (3B activated), with 8bit quantization, but I haven't been > willing to trust it with generating kernel code. I have experimenting > to see how it compares with Gemini 3.1 when creating a python script > to send e-mail[1] or creating a bash completion script[2] for my > fstests test appliance. Most of my tests are also to generate ancillary python scripts too. I did some tests using it and using chatgpt and deepseek (using a free account). On my tests, qwen3.6 code (36B, 3B activated) had similar quality. As we're aiming on patch review, we don't need a model capable of generating kernel code. Instead, we're aiming on one that helps to review it. A well trained model using lore may end giving similar results. As a quality criteria, I'd say if the simpler model would have maybe 80%-90% quality when compared with using a frontier model, it sounds worth having it, not only for people to run locally but also as a continuity plan if/when we end running out of free passes to run it on server-grade GPU hardware. > [1] https://github.com/tytso/xfstests-bld/commit/dfadb2014da446ecb967de51904ee531f7be8bd5 > [2] https://github.com/tytso/xfstests-bld/commit/dfadb2014da446ecb967de51904ee531f7be8bd5 > > However, from Roman tells me, Sashiko is running muliple LLM passes > using a frontier model for each commit review. So what Sashiko does > is quite a bit more complicated than a series of prompts such as: > > Create a python program which submits an e-mail message using the > Submission Port (port 587), It should enable encryption using > STARTTLS and it should obtain the username and password from a > config.ini file. Model the python program using the send-mail.py > in the sandbox directory. It should support the same command-line > options but instead of sending the e-mail using sendgrid, it should > send the e-mail using the Submission protocol. > > Please add pydoc documentation to send-mail-smtp.py. > > Please enhance the program you just created (send-mail-smtp.py) to > support specifying a path to a certificate file in the > configuration file in case the user doesn't want to use the system > provided top-level trusted certificates. > > Please add support for a configuration file parameter which > specifies whether TLS should be mandatory, optional, or disabled. > Update the pydoc documentation as necessary. > > Please integrate the functionality found in send-mail.py into > send-mail-smtp.py. Instead of getting the Sendgrid API key from an > environment variable, change it to obtain the Sendgrid API key from > the configuration file. If the Sendgrid API key is specified, use > sendgrid instead of the SMTP submission protocol. Update the pydoc > documentation strings. > > .... which does work pretty well even on less capable models that I > can run locally. > > I guess we could try running Sashiko using ollama-mlx on a Macbook > with 128GB, and see how it works, but my assumption is that the answer > is "not well" --- which is why I really want to look at fine-tuning > one of these smaller models first. I agree with you: without distilling a smaller model, it may not work well. However, frontier models may spend a lot more tokens than small models for the same prompt, due to mult-step reasoning and the huge amount of parameters, so we could have some surprise here. Sounds worth trying it. Thanks, Mauro