Re: Slow down when running newer Cairo on ARM with NEON

Bill Spitzak <[email protected]> Tue, 16 Apr 2019 15:37:32 -0700
Newsgroups gmane.comp.lib.cairo
Message-ID <CAL-8oAjdyTQY2+wk2-iDASsc3OYUmo5XactOpCSHYzxwjyc2Cg@mail.gmail.com>
--===============1370199926==
Content-Type: multipart/alternative; boundary="000000000000cf5f6f0586ad6af6"

--000000000000cf5f6f0586ad6af6
Content-Type: text/plain; charset="UTF-8"

On Tue, Apr 16, 2019 at 1:53 PM Joshua Watt <[email protected]> wrote:

> On Tue, 2019-04-16 at 12:01 -0700, Bill Spitzak wrote:
>
> I think you can force the interpolation to bilinear or impulse. However
> you are going to revert to 1980's style scaling with extreme aliasing.
>
>
> Hmm, this seems to be empirically false (at least when using ARM+NEON and
> pixman 0.34). If I revert the cairo change GOOD quality images both look
> fine and render quickly. If I change my application to use FAST quality,
> then I certainly see the 1980's graphics with aliasing and it is even
> faster still, but I'd rather find a way to keep the good quality images
> without the 70% performance hit.
>

Try BILINEAR which should be what GOOD was doing before.


>
> This is some of my work from 4 years ago and unfortunately it never got
> finished due to rejection by the Pixman maintainers (who I think may not be
> working on it any more). This was to implement a two-pass algorithm in
> Pixman that could also do non-affine (perspective) transforms. It should be
> considerably faster for any down-scaling, even if the filter is set to
> bilinear. The current code is not 2-pass (in effect both passes are run for
> every output pixel, rather than saving the result of the first pass, this
> is in fact worse than convolving with a 2-D filter), but at least produces
> modern results.
>
> The problem is that the filters cannot be specified as arrays of weights,
> due to the need to choose arbitrary filter sizes, both to allow non-affine
> transforms and just because most 2-pass algorithms require unexpected
> filter sizes (such as the derivative along the x axis of the input but the
> y axis of the output). IMHO the most practical way to get this is to just
> make "GOOD" and "BEST" select two implementation-chosen filters (BILINEAR
> and IMPULSE would also be allowed) and scrap any ability to specify the
> filter more accurately by the client. This seemed to produce considerable
> pushback in pixman and was rejected and I gave up after succeeding in
> getting the api implemented in Cairo.
>
>
> Perhaps the NEON implementation in pixman is doing something more modern
> and thus is fast and good quality? I will admit that image interpolation
> algorithms aren't my area of expertise and I don't really follow the
> details of what you are saying here, I'm just reporting what I see
> empirically :)
>
>
> On Tue, Apr 16, 2019 at 10:38 AM Joshua Watt <[email protected]> wrote:
>
> Hello,
>
> I recently upgrade from Cairo 1.12 to 1.14 (yes, I know these are old
> versions), and after doing so noticed a approximately 70% reduction in
> performance when rendering scenes that make heavy use of image scaling.
> I did some digging and tracking the offending commit down to the
> commit: f337342c8 ("V6 image: Use convolution filters for sample
> reconstruction when downscaling")
>
> It appears that this commit is attempting to improve the quality of
> downscaled images by implementing new interpolation algorithms in cairo
> instead of using the pixman algorithms. My theory is that this is much
> slower on ARM processes that have NEON support because pixman has
> special implementations of the interpolations algorithms written to
> take advantage of NEON, while the new cairo implementations do not.
>
> Does anyone have any ideas on what a good path forward would be to
> restore the ARM+NEON performance? I am planning on trying to reproduce
> this with a newer version of cairo to see if it is still a problem, but
> I suspect it will be based on the lack of any significant changes in
> this code to either cairo or pixman.
>
> --
> Joshua Watt <[email protected]>
>
> --
>
> Joshua Watt <[email protected]>
>

--000000000000cf5f6f0586ad6af6
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><div dir=3D"ltr"><br></div><br><div class=3D"gmail_quote">=
<div dir=3D"ltr" class=3D"gmail_attr">On Tue, Apr 16, 2019 at 1:53 PM Joshu=
a Watt &lt;<a href=3D"mailto:[email protected]">[email protected]</a>=
&gt; wrote:<br></div><blockquote class=3D"gmail_quote" style=3D"margin:0px =
0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"><div=
 style=3D"text-align:left;direction:ltr"><div>On Tue, 2019-04-16 at 12:01 -=
0700, Bill Spitzak wrote:</div><blockquote type=3D"cite" style=3D"margin:0p=
x 0px 0px 0.8ex;border-left:2px solid rgb(114,159,207);padding-left:1ex"><d=
iv dir=3D"ltr"><div dir=3D"ltr">I think you can force the interpolation to =
bilinear or impulse. However you are going to revert to 1980&#39;s style sc=
aling with extreme aliasing.</div></div></blockquote><div><br></div><div>Hm=
m, this seems to be empirically false (at least when using ARM+NEON and pix=
man 0.34). If I revert the cairo change GOOD quality images both look fine =
and render quickly. If I change my application to use FAST quality, then I =
certainly see the 1980&#39;s graphics with aliasing and it is even faster s=
till, but I&#39;d rather find a way to keep the good quality images without=
 the 70% performance hit.</div></div></blockquote><div><br></div><div>Try B=
ILINEAR which should be what GOOD was doing before.</div><div><br></div><bl=
ockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-lef=
t:1px solid rgb(204,204,204);padding-left:1ex"><div style=3D"text-align:lef=
t;direction:ltr"><div><br></div><blockquote type=3D"cite" style=3D"margin:0=
px 0px 0px 0.8ex;border-left:2px solid rgb(114,159,207);padding-left:1ex"><=
div dir=3D"ltr"><div dir=3D"ltr"><div><br></div><div>This is some of my wor=
k from 4 years ago and unfortunately it never got finished due to rejection=
 by the Pixman maintainers (who I think may not be working on it any more).=
 This was to implement a two-pass algorithm in Pixman that could also do no=
n-affine (perspective) transforms. It should be considerably faster for any=
 down-scaling, even if the filter is set to bilinear. The current code is n=
ot 2-pass (in effect both passes are run for every output pixel, rather tha=
n saving the result of the first pass, this is in fact worse than convolvin=
g with a 2-D filter), but at least produces modern results.</div><div><br><=
/div><div>The problem is that the filters cannot be specified as arrays of =
weights, due to the need to choose arbitrary filter sizes, both to allow no=
n-affine transforms and just because most 2-pass algorithms require unexpec=
ted filter sizes (such as the derivative along the x axis of the input but =
the y axis of the output). IMHO the most practical way to get this is to ju=
st make &quot;GOOD&quot; and &quot;BEST&quot; select two implementation-cho=
sen filters (BILINEAR and IMPULSE would also be allowed) and scrap any abil=
ity to specify the filter more accurately by the client. This seemed to pro=
duce considerable pushback in pixman and was rejected and I gave up after s=
ucceeding in getting the api implemented in Cairo.</div><div><br></div></di=
v></div></blockquote><div><br></div><div><div>Perhaps the NEON implementati=
on in pixman is doing something more modern and thus is fast and good quali=
ty? I will admit that image interpolation algorithms aren&#39;t my area of =
expertise and I don&#39;t really follow the details of what you are saying =
here, I&#39;m just reporting what I see empirically :)</div><div><br></div>=
<blockquote type=3D"cite" style=3D"margin:0px 0px 0px 0.8ex;border-left:2px=
 solid rgb(114,159,207);padding-left:1ex"><div dir=3D"ltr"><div dir=3D"ltr"=
></div></div></blockquote></div><blockquote type=3D"cite" style=3D"margin:0=
px 0px 0px 0.8ex;border-left:2px solid rgb(114,159,207);padding-left:1ex"><=
br><div class=3D"gmail_quote"><div dir=3D"ltr" class=3D"gmail_attr">On Tue,=
 Apr 16, 2019 at 10:38 AM Joshua Watt &lt;<a href=3D"mailto:jpewhacker@gmai=
l.com" target=3D"_blank">[email protected]</a>&gt; wrote:<br></div><bloc=
kquote type=3D"cite" style=3D"margin:0px 0px 0px 0.8ex;border-left:2px soli=
d rgb(114,159,207);padding-left:1ex">Hello,<br>
<br>
I recently upgrade from Cairo 1.12 to 1.14 (yes, I know these are old<br>
versions), and after doing so noticed a approximately 70% reduction in<br>
performance when rendering scenes that make heavy use of image scaling.<br>
I did some digging and tracking the offending commit down to the<br>
commit: f337342c8 (&quot;V6 image: Use convolution filters for sample<br>
reconstruction when downscaling&quot;)<br>
<br>
It appears that this commit is attempting to improve the quality of<br>
downscaled images by implementing new interpolation algorithms in cairo<br>
instead of using the pixman algorithms. My theory is that this is much<br>
slower on ARM processes that have NEON support because pixman has<br>
special implementations of the interpolations algorithms written to<br>
take advantage of NEON, while the new cairo implementations do not.<br>
<br>
Does anyone have any ideas on what a good path forward would be to<br>
restore the ARM+NEON performance? I am planning on trying to reproduce<br>
this with a newer version of cairo to see if it is still a problem, but<br>
I suspect it will be based on the lack of any significant changes in<br>
this code to either cairo or pixman.<br>
<br>
-- <br>
Joshua Watt &lt;<a href=3D"mailto:[email protected]" target=3D"_blank">J=
[email protected]</a>&gt;<br>
<br>
</blockquote></div></blockquote><div><span><pre>-- <br></pre>Joshua Watt &l=
t;<a href=3D"mailto:[email protected]" target=3D"_blank">JPEWhacker@gmai=
l.com</a>&gt;</span></div></div>
</blockquote></div></div>

--000000000000cf5f6f0586ad6af6--

--===============1370199926==
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: inline

LS0gCmNhaXJvIG1haWxpbmcgbGlzdApjYWlyb0BjYWlyb2dyYXBoaWNzLm9yZwpodHRwczovL2xp
c3RzLmNhaXJvZ3JhcGhpY3Mub3JnL21haWxtYW4vbGlzdGluZm8vY2Fpcm8=

--===============1370199926==--