Re: Better compatibility of the Python scientific/data stack with fast Python interpreters

Ralf Gommers via NumPy-Discussion <[email protected]> Wed, 30 Apr 2025 07:32:44 +0200
Newsgroups gmane.comp.python.numeric.general,gmane.comp.python.hpy,gmane.comp.python.pypy
Message-ID <CABL7CQgzTpEKgb5wGWy=kk6y2fzt7yPBmuckjzXkhnB740UM2w@mail.gmail.com>
--===============2092502468507426316==
Content-Type: multipart/alternative; boundary="000000000000cbb2f30633f83f9d"

--000000000000cbb2f30633f83f9d
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

On Tue, Apr 29, 2025 at 11:24=E2=80=AFAM PIERRE AUGIER <
[email protected]> wrote:

> Dear Numpy community members and Numpy developers,
>
> This email is to get the points of view of the Numpy community members an=
d
> developers about a subject that I find very important. I'm going to
> introduce it just in few lines so I write few statements without backing
> them with proper arguments and without giving links and word definitions.=
 I
> assume that most people registered to the numpy-discussion list are at
> least familiar to the subject.
>

Thanks for thinking about this Pierre.


> I think getting proper compatibility of the Python scientific/data stack
> with fast Python interpreters is very important for the long-term future =
of
> Python for science and data.
>

I'm not sure it is, it wouldn't rank high on my wish list. PyPy is nearing
its end-of-life, and GraalPy is really only for Java devs it seems - I've
never seen it in the real world. With CPython performance being worked hard
on, the performance gap is shrinking every year.

More importantly, none of these efforts (including the "faster CPython"
project), seem critical to numerical/scientific users. We're still talking
about pure Python code that gets maybe up to 5x faster, while the gains
from doing things in compiled languages are a lot higher. So the benefits
are more important for small packages if it moves the threshold at which it
becomes necessary for them to write zero extension modules. For core
libraries like NumPy, pure Python performance isn't super critical.

When thinking about overall performance improvements, I'd say that the two
most promising large developments are: (1) making it easier to use
accelerator libraries (PyTorch, CuPy et al.), and (2) free-threaded CPython
for enabling Python-level threading.


> Nowadays, fast Python implementations like PyPy and GraalPy are
>
> - complicated to use (need local builds of wheels)
> - slow as soon as the code involves Python extensions (because extensions
> work through an emulation layer, like cpyext in PyPy)
>
> Therefore, these fast Python implementations are not as popular as they
> should be. However, they are really great tools which should be used in
> particular for scientific/data applications.
>
> Unfortunately, if we just follow the pace of the CPython C API evolution,
> we will basically never have an ecosystem natively compatible with fast
> Python implementations.
>
> The more I think and learn about this subject, the more I think that Nump=
y
> has to stop using directly the CPython C API and to be ported to HPy
> (though there are other alternatives - in particular Nanobind and Cython =
-
> that could be discussed). Numpy 1 has been ported to HPy but unfortunatel=
y,
> there have been deep changes with Numpy 2 (and Meson) so it seems that on=
e
> should restart the porting.
>
> Moreover, unfortunately, HPy does not currently receive as much care as i=
t
> should.
>

> It seems to me that the project of fixing the roots of the Python
> ecosystem has to be relaunched. I think that the dynamics has to come fro=
m
> the Python scientific/data community and in particular Numpy. It is
> unfortunately outside of the C API working group's scope (see
> https://discuss.python.org/t/c-api-working-group-and-plan-to-get-a-python=
-c-api-compatible-with-alternative-python-implementations/89477
> ).
>
> It seems to me that it is necessary (and possible) to get some founding
> for such an impactful project so that we can get people working on it.
>
> I'd like to write a long and serious text on this subject, and I first tr=
y
> to get the points of view of the different people and projects involved.
>
> I guess I should write explicit questions:
>
> - What do you think about the project of fixing the Python scientific/dat=
a
> stack so that it becomes natively compatible (hence fast and convenient,
> with easy installations) with alternative and fast Python interpreters?
> - Do you have points of view on how this should be done, technically
> (HPy?, something else?) and on other aspects (community, NEP?, founding,
> ...).
> - Anything else interesting on this subject?
>

Having HPy or something like it will be very nice and lead ot long-term
benefits. The problem with HPy seems to be more a social one at this point:
if CPython core devs don't want to adopt it but do their own "make the C
API more opaque" strategy, then more effort on HPy isn't going to help. If
you're going to dig into this more, I suggest trying to get a very good
sense of what the CPython core dev team, and in particular its C API
Workgroup, is thinking/planning. That will inform whether the right
strategy is to help their efforts along, or work on HPy.

Cheers,
Ralf



>
> Best regards,
> Pierre
>
> --
> Pierre Augier - CR CNRS                 http://www.legi.grenoble-inp.fr
> LEGI (UMR 5519) Laboratoire des Ecoulements Geophysiques et Industriels
> BP53, 38041 Grenoble Cedex, France                tel:+33.4.56.52.86.16
> _______________________________________________
> NumPy-Discussion mailing list -- [email protected]
> To unsubscribe send an email to [email protected]
> https://mail.python.org/mailman3/lists/numpy-discussion.python.org/
> Member address: [email protected]
>

--000000000000cbb2f30633f83f9d
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><div class=3D"gmail_quote gmail_quote_container"><div =
dir=3D"ltr" class=3D"gmail_attr">On Tue, Apr 29, 2025 at 11:24=E2=80=AFAM P=
IERRE AUGIER &lt;<a href=3D"mailto:[email protected]">pi=
[email protected]</a>&gt; wrote:<br></div><blockquote clas=
s=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid r=
gb(204,204,204);padding-left:1ex">Dear Numpy community members and Numpy de=
velopers,<br>
<br>
This email is to get the points of view of the Numpy community members and =
developers about a subject that I find very important. I&#39;m going to int=
roduce it just in few lines so I write few statements without backing them =
with proper arguments and without giving links and word definitions. I assu=
me that most people registered to the numpy-discussion list are at least fa=
miliar to the subject.<br></blockquote><div><br></div><div>Thanks for think=
ing about this Pierre. <br></div><div>=C2=A0<br></div><blockquote class=3D"=
gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(20=
4,204,204);padding-left:1ex">
I think getting proper compatibility of the Python scientific/data stack wi=
th fast Python interpreters is very important for the long-term future of P=
ython for science and data.<br></blockquote><div><br></div><div>I&#39;m not=
 sure it is, it wouldn&#39;t rank high on my wish list. PyPy is nearing its=
 end-of-life, and GraalPy is really only for Java devs it seems - I&#39;ve =
never seen it in the real world. With CPython performance being worked hard=
 on, the performance gap is shrinking every year.</div><div><br></div><div>=
More importantly, none of these efforts (including the &quot;faster CPython=
&quot; project), seem critical to numerical/scientific users. We&#39;re sti=
ll talking about pure Python code that gets maybe up to 5x faster, while th=
e gains from doing things in compiled languages are a lot higher. So the be=
nefits are more important for small packages if it moves the threshold at w=
hich it becomes necessary for them to write zero extension modules. For cor=
e libraries like NumPy, pure Python performance isn&#39;t super critical.</=
div><div><br></div><div>When thinking about overall performance improvement=
s, I&#39;d say that the two most promising large developments are: (1) maki=
ng it easier to use accelerator libraries (PyTorch, CuPy et al.), and (2) f=
ree-threaded CPython for enabling Python-level threading. <br></div><div><b=
r> </div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8e=
x;border-left:1px solid rgb(204,204,204);padding-left:1ex">
<br>
Nowadays, fast Python implementations like PyPy and GraalPy are <br>
<br>
- complicated to use (need local builds of wheels)<br>
- slow as soon as the code involves Python extensions (because extensions w=
ork through an emulation layer, like cpyext in PyPy)<br>
<br>
Therefore, these fast Python implementations are not as popular as they sho=
uld be. However, they are really great tools which should be used in partic=
ular for scientific/data applications.<br>
<br>
Unfortunately, if we just follow the pace of the CPython C API evolution, w=
e will basically never have an ecosystem natively compatible with fast Pyth=
on implementations.<br>
<br>
The more I think and learn about this subject, the more I think that Numpy =
has to stop using directly the CPython C API and to be ported to HPy (thoug=
h there are other alternatives - in particular Nanobind and Cython - that c=
ould be discussed). Numpy 1 has been ported to HPy but unfortunately, there=
 have been deep changes with Numpy 2 (and Meson) so it seems that one shoul=
d restart the porting.<br>
<br>
Moreover, unfortunately, HPy does not currently receive as much care as it =
should. <br></blockquote><blockquote class=3D"gmail_quote" style=3D"margin:=
0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex">
<br>
It seems to me that the project of fixing the roots of the Python ecosystem=
 has to be relaunched. I think that the dynamics has to come from the Pytho=
n scientific/data community and in particular Numpy. It is unfortunately ou=
tside of the C API working group&#39;s scope (see <a href=3D"https://discus=
s.python.org/t/c-api-working-group-and-plan-to-get-a-python-c-api-compatibl=
e-with-alternative-python-implementations/89477" rel=3D"noreferrer" target=
=3D"_blank">https://discuss.python.org/t/c-api-working-group-and-plan-to-ge=
t-a-python-c-api-compatible-with-alternative-python-implementations/89477</=
a>).<br>
<br>
It seems to me that it is necessary (and possible) to get some founding for=
 such an impactful project so that we can get people working on it.<br>
<br>
I&#39;d like to write a long and serious text on this subject, and I first =
try to get the points of view of the different people and projects involved=
.<br>
<br>
I guess I should write explicit questions:<br>
<br>
- What do you think about the project of fixing the Python scientific/data =
stack so that it becomes natively compatible (hence fast and convenient, wi=
th easy installations) with alternative and fast Python interpreters?<br>
- Do you have points of view on how this should be done, technically (HPy?,=
 something else?) and on other aspects (community, NEP?, founding, ...).<br=
>
- Anything else interesting on this subject?<br></blockquote><div><br></div=
><div><div>Having HPy or something like it will be very nice and lead ot lo=
ng-term benefits. The problem with HPy seems to be more a social one at thi=
s point: if CPython core devs don&#39;t want to adopt it but do their own &=
quot;make the C API more opaque&quot; strategy, then more effort on HPy isn=
&#39;t going to help. If you&#39;re going to dig into this more, I suggest =
trying to get a very good sense of what the CPython core dev team, and in p=
articular its C API Workgroup, is thinking/planning. That will inform wheth=
er the right strategy is to help their efforts along, or work on HPy.</div>=
<div><br></div><div>Cheers,</div><div>Ralf</div><div><br></div>=C2=A0</div>=
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left:1px solid rgb(204,204,204);padding-left:1ex">
<br>
Best regards,<br>
Pierre<br>
<br>
--<br>
Pierre Augier - CR CNRS=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0 =C2=A0<a href=3D"http://www.legi.grenoble-inp.fr" rel=3D"noreferrer" ta=
rget=3D"_blank">http://www.legi.grenoble-inp.fr</a><br>
LEGI (UMR 5519) Laboratoire des Ecoulements Geophysiques et Industriels<br>
BP53, 38041 Grenoble Cedex, France=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0=
 =C2=A0 =C2=A0 tel:+33.4.56.52.86.16<br>
_______________________________________________<br>
NumPy-Discussion mailing list -- <a href=3D"mailto:numpy-discussion@python.=
org" target=3D"_blank">[email protected]</a><br>
To unsubscribe send an email to <a href=3D"mailto:numpy-discussion-leave@py=
thon.org" target=3D"_blank">[email protected]</a><br>
<a href=3D"https://mail.python.org/mailman3/lists/numpy-discussion.python.o=
rg/" rel=3D"noreferrer" target=3D"_blank">https://mail.python.org/mailman3/=
lists/numpy-discussion.python.org/</a><br>
Member address: <a href=3D"mailto:[email protected]" target=3D"_b=
lank">[email protected]</a><br>
</blockquote></div></div>

--000000000000cbb2f30633f83f9d--

--===============2092502468507426316==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
NumPy-Discussion mailing list -- [email protected]
To unsubscribe send an email to [email protected]
https://mail.python.org/mailman3/lists/numpy-discussion.python.org/
Member address: [email protected]

--===============2092502468507426316==--