Re: Proposal -- Interpretation of DFSG on Artificial Intelligence (AI) Models
Soren Stoutner <[email protected]>
| Newsgroups | gmane.linux.debian.devel.vote |
|---|---|
| Organization | Debian |
| Message-ID | <11728906.MucGe3eQFb@soren-desktop> |
On Wednesday, May 14, 2025 5:04:03 PM Mountain Standard Time Arian Ott wrote: > During the course of my semester thesis on Retrieval-Augmented Generation > (RAG), I encountered a compelling example wherein an AI model identified a > previously unknown biomarker associated with cancer. This discovery was > only possible because the researchers had access to the underlying dataset. > Without that access, the model’s findings would have been opaque and > potentially unverifiable. > > This brings me to a central concern: when data scientists are given a model > to work with, their first question is often: > “What data was used to train it?” > This question is not incidental. It is fundamental to understanding the > model’s behaviour, biases, and limitations. It is also essential for > scientific reproducibility. That is a good, concrete example. It is interesting that access to the original training data has value that goes beyond a desire to retrain the model and extends into *using* the model to its fullest extent. > In the course of the earlier email exchange, it was argued that the > hardware requirements for training large-scale models place them out of > reach for anyone without a budget in the range of 100 M€. While this may be > true for frontier-scale models, I believe it overlooks a significant > portion of real-world use cases. > > In my undergraduate work, we frequently relied on publicly available > datasets from sources such as Kaggle. These enabled us to train our own > models, interpret results, and explore data-driven questions in a hands-on > manner. Providing access to training data empowers researchers, > institutions, and independent developers to create models adapted to their > specific needs. Moreover, it facilitates the composability of data, an > essential feature in interdisciplinary research and real-world applications. Out of curiosity, how much hardware did you need to train your own models on these data sets? I think sometimes we forget that many of the MLs use data sets that are much smaller than LLMs scraping the entire web. -- Soren Stoutner [email protected]
signature.asc
(application/pgp-signature, 833 B)
-----BEGIN PGP SIGNATURE----- iQIzBAABCgAdFiEEJKVN2yNUZnlcqOI+wufLJ66wtgMFAmglM9YACgkQwufLJ66w tgPDlRAAp3CpyL1/lBgFmkBulTkzI9xEfPPAmLaTfdVCiz4tIaiA5Io1R/OjAlBW RUlcspO0d36MI2I/Gxh4az7FGOAURbaC5Nw0sNYWP+ibSBJ/P4Qa8LwCxBPykAr4 d+5O8/q50J508PmQ2QABmcBUNPA0dRcBOrlElroNup7DOl4N51OxgPxU+7zxuk67 AfbZF66B6imTDCEQpPwElAdgK4HO3fV6M4dND1YpLVr5nEuK8vIgxIWl2uVuiDBE k0Lig4rll1OG+/myu99QgL/9R1XhTkHatm8WR61S1nm/taHDiM/JbrAGJEbzo9/W N308lN5DLH0gZmIjOasdMooZ98WWtPzGkgcsox8lhrk8LqmRXOU0vUiVzc36U4cx uu9cESHlfU2tDqveY/Pbb2e3+FZktjvSh4CKRaT54pvWH/poCuoLRw/szh/AxOOR byHxnApVDc3NGHjTHfTPIO1IBdngxvQ4BYF/IY5tLeg4ABjxyiagoo1tWvXKySk+ H28jCFLEWXA7vaLiUxJPcpWBjfS12ak7QNLh1bW+cV/+v1d+EEG6C6WtUPPjisbY r3A0HYXVmG26GA44QbVk0uyvW61Mh5mX8Zc8mUxhTDaUsm4+OmR0K+RE5ylkmxEg D36xTOY6+xTWgushFEQWViJMYN/TG9Eb0y9t5HsttoY/Uw8lI1A= =hwOQ -----END PGP SIGNATURE-----