Re: Proposal -- Interpretation of DFSG on Artificial Intelligence (AI) Models

Soren Stoutner <[email protected]>
Newsgroups gmane.linux.debian.devel.vote
Organization Debian
Message-ID <11728906.MucGe3eQFb@soren-desktop>
On Wednesday, May 14, 2025 5:04:03 PM Mountain Standard Time Arian Ott wrote:
> During the course of my semester thesis on Retrieval-Augmented Generation
> (RAG), I encountered a compelling example wherein an AI model identified a
> previously unknown biomarker associated with cancer. This discovery was
> only possible because the researchers had access to the underlying dataset.
> Without that access, the model’s findings would have been opaque and
> potentially unverifiable.
> 
> This brings me to a central concern: when data scientists are given a model
> to work with, their first question is often:
> “What data was used to train it?”
> This question is not incidental. It is fundamental to understanding the
> model’s behaviour, biases, and limitations. It is also essential for
> scientific reproducibility.

That is a good, concrete example.  It is interesting that access to the 
original training data has value that goes beyond a desire to retrain the 
model and extends into *using* the model to its fullest extent.
 
> In the course of the earlier email exchange, it was argued that the
> hardware requirements for training large-scale models place them out of
> reach for anyone without a budget in the range of 100 M€. While this may be
> true for frontier-scale models, I believe it overlooks a significant
> portion of real-world use cases.
> 
> In my undergraduate work, we frequently relied on publicly available
> datasets from sources such as Kaggle. These enabled us to train our own
> models, interpret results, and explore data-driven questions in a hands-on
> manner. Providing access to training data empowers researchers,
> institutions, and independent developers to create models adapted to their
> specific needs. Moreover, it facilitates the composability of data, an
> essential feature in interdisciplinary research and real-world applications.

Out of curiosity, how much hardware did you need to train your own models on 
these data sets?  I think sometimes we forget that many of the MLs use data 
sets that are much smaller than LLMs scraping the entire web.

-- 
Soren Stoutner
[email protected]
signature.asc (application/pgp-signature, 833 B)
-----BEGIN PGP SIGNATURE-----

iQIzBAABCgAdFiEEJKVN2yNUZnlcqOI+wufLJ66wtgMFAmglM9YACgkQwufLJ66w
tgPDlRAAp3CpyL1/lBgFmkBulTkzI9xEfPPAmLaTfdVCiz4tIaiA5Io1R/OjAlBW
RUlcspO0d36MI2I/Gxh4az7FGOAURbaC5Nw0sNYWP+ibSBJ/P4Qa8LwCxBPykAr4
d+5O8/q50J508PmQ2QABmcBUNPA0dRcBOrlElroNup7DOl4N51OxgPxU+7zxuk67
AfbZF66B6imTDCEQpPwElAdgK4HO3fV6M4dND1YpLVr5nEuK8vIgxIWl2uVuiDBE
k0Lig4rll1OG+/myu99QgL/9R1XhTkHatm8WR61S1nm/taHDiM/JbrAGJEbzo9/W
N308lN5DLH0gZmIjOasdMooZ98WWtPzGkgcsox8lhrk8LqmRXOU0vUiVzc36U4cx
uu9cESHlfU2tDqveY/Pbb2e3+FZktjvSh4CKRaT54pvWH/poCuoLRw/szh/AxOOR
byHxnApVDc3NGHjTHfTPIO1IBdngxvQ4BYF/IY5tLeg4ABjxyiagoo1tWvXKySk+
H28jCFLEWXA7vaLiUxJPcpWBjfS12ak7QNLh1bW+cV/+v1d+EEG6C6WtUPPjisbY
r3A0HYXVmG26GA44QbVk0uyvW61Mh5mX8Zc8mUxhTDaUsm4+OmR0K+RE5ylkmxEg
D36xTOY6+xTWgushFEQWViJMYN/TG9Eb0y9t5HsttoY/Uw8lI1A=
=hwOQ
-----END PGP SIGNATURE-----
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.