Re: Slow down when running newer Cairo on ARM with NEON

Joshua Watt <[email protected]> Tue, 16 Apr 2019 15:53:44 -0500
Newsgroups gmane.comp.lib.cairo
Message-ID <[email protected]>
--===============1078554132==
Content-Type: multipart/alternative; boundary="=-Ugj1gjdK8NLHmuQJAYBx"


--=-Ugj1gjdK8NLHmuQJAYBx
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: 7bit

On Tue, 2019-04-16 at 12:01 -0700, Bill Spitzak wrote:
> I think you can force the interpolation to bilinear or impulse.
> However you are going to revert to 1980's style scaling with extreme
> aliasing.

Hmm, this seems to be empirically false (at least when using ARM+NEON
and pixman 0.34). If I revert the cairo change GOOD quality images both
look fine and render quickly. If I change my application to use FAST
quality, then I certainly see the 1980's graphics with aliasing and it
is even faster still, but I'd rather find a way to keep the good
quality images without the 70% performance hit.
> This is some of my work from 4 years ago and unfortunately it never
> got finished due to rejection by the Pixman maintainers (who I think
> may not be working on it any more). This was to implement a two-pass
> algorithm in Pixman that could also do non-affine (perspective)
> transforms. It should be considerably faster for any down-scaling,
> even if the filter is set to bilinear. The current code is not 2-pass 
> (in effect both passes are run for every output pixel, rather than
> saving the result of the first pass, this is in fact worse than
> convolving with a 2-D filter), but at least produces modern results.
> 
> The problem is that the filters cannot be specified as arrays of
> weights, due to the need to choose arbitrary filter sizes, both to
> allow non-affine transforms and just because most 2-pass algorithms
> require unexpected filter sizes (such as the derivative along the x
> axis of the input but the y axis of the output). IMHO the most
> practical way to get this is to just make "GOOD" and "BEST" select
> two implementation-chosen filters (BILINEAR and IMPULSE would also be
> allowed) and scrap any ability to specify the filter more accurately
> by the client. This seemed to produce considerable pushback in pixman
> and was rejected and I gave up after succeeding in getting the api
> implemented in Cairo.
> 

Perhaps the NEON implementation in pixman is doing something more
modern and thus is fast and good quality? I will admit that image
interpolation algorithms aren't my area of expertise and I don't really
follow the details of what you are saying here, I'm just reporting what
I see empirically :)
> On Tue, Apr 16, 2019 at 10:38 AM Joshua Watt <[email protected]>
> wrote:
> > Hello,
> > 
> > 
> > 
> > I recently upgrade from Cairo 1.12 to 1.14 (yes, I know these are
> > old
> > 
> > versions), and after doing so noticed a approximately 70% reduction
> > in
> > 
> > performance when rendering scenes that make heavy use of image
> > scaling.
> > 
> > I did some digging and tracking the offending commit down to the
> > 
> > commit: f337342c8 ("V6 image: Use convolution filters for sample
> > 
> > reconstruction when downscaling")
> > 
> > 
> > 
> > It appears that this commit is attempting to improve the quality of
> > 
> > downscaled images by implementing new interpolation algorithms in
> > cairo
> > 
> > instead of using the pixman algorithms. My theory is that this is
> > much
> > 
> > slower on ARM processes that have NEON support because pixman has
> > 
> > special implementations of the interpolations algorithms written to
> > 
> > take advantage of NEON, while the new cairo implementations do not.
> > 
> > 
> > 
> > Does anyone have any ideas on what a good path forward would be to
> > 
> > restore the ARM+NEON performance? I am planning on trying to
> > reproduce
> > 
> > this with a newer version of cairo to see if it is still a problem,
> > but
> > 
> > I suspect it will be based on the lack of any significant changes
> > in
> > 
> > this code to either cairo or pixman.
> > 
> > 
> > 
> > -- 
> > 
> > Joshua Watt <[email protected]>
> > 
> > 
> > 
-- 
Joshua Watt <[email protected]>

--=-Ugj1gjdK8NLHmuQJAYBx
Content-Type: text/html; charset="utf-8"
Content-Transfer-Encoding: quoted-printable

<html dir=3D"ltr"><head></head><body style=3D"text-align:left; direction:lt=
r;"><div>On Tue, 2019-04-16 at 12:01 -0700, Bill Spitzak wrote:</div><block=
quote type=3D"cite" style=3D"margin:0 0 0 .8ex; border-left:2px #729fcf sol=
id;padding-left:1ex"><div dir=3D"ltr"><div dir=3D"ltr">I think you can forc=
e the interpolation to bilinear or impulse. However you are going to revert=
 to 1980's style scaling with extreme aliasing.</div></div></blockquote><di=
v><br></div><div>Hmm, this seems to be empirically false (at least when usi=
ng ARM+NEON and pixman 0.34). If I revert the cairo change GOOD quality ima=
ges both look fine and render quickly. If I change my application to use FA=
ST quality, then I certainly see the 1980's graphics with aliasing and it i=
s even faster still, but I'd rather find a way to keep the good quality ima=
ges without the 70% performance hit.</div><div><br></div><blockquote type=
=3D"cite" style=3D"margin:0 0 0 .8ex; border-left:2px #729fcf solid;padding=
-left:1ex"><div dir=3D"ltr"><div dir=3D"ltr"><div><br></div><div>This is so=
me of my work from 4 years ago and unfortunately it never got finished due =
to rejection by the Pixman maintainers (who I think may not be working on i=
t any more). This was to implement a two-pass algorithm in Pixman that coul=
d also do non-affine (perspective) transforms. It should be considerably fa=
ster for any down-scaling, even if the filter is set to bilinear. The curre=
nt code is not 2-pass (in effect both passes are run for every output pixel=
, rather than saving the result of the first pass, this is in fact worse th=
an convolving with a 2-D filter), but at least produces modern results.</di=
v><div><br></div><div>The problem is that the filters cannot be specified a=
s arrays of weights, due to the need to choose arbitrary filter sizes, both=
 to allow non-affine transforms and just because most 2-pass algorithms req=
uire unexpected filter sizes (such as the derivative along the x axis of th=
e input but the y axis of the output). IMHO the most practical way to get t=
his is to just make "GOOD" and "BEST" select two implementation-chosen filt=
ers (BILINEAR and IMPULSE would also be allowed) and scrap any ability to s=
pecify the filter more accurately by the client. This seemed to produce con=
siderable pushback in pixman and was rejected and I gave up after succeedin=
g in getting the api implemented in Cairo.</div><div><br></div></div></div>=
</blockquote><div><br></div><div><div>Perhaps the NEON implementation in pi=
xman is doing something more modern and thus is fast and good quality? I wi=
ll admit that image interpolation algorithms aren't my area of expertise an=
d I don't really follow the details of what you are saying here, I'm just r=
eporting what I see empirically :)</div><div><br></div><blockquote type=3D"=
cite" style=3D"margin:0 0 0 .8ex; border-left:2px #729fcf solid;padding-lef=
t:1ex"><div dir=3D"ltr"><div dir=3D"ltr"></div></div></blockquote></div><bl=
ockquote type=3D"cite" style=3D"margin:0 0 0 .8ex; border-left:2px #729fcf =
solid;padding-left:1ex"><br><div class=3D"gmail_quote"><div dir=3D"ltr" cla=
ss=3D"gmail_attr">On Tue, Apr 16, 2019 at 10:38 AM Joshua Watt &lt;<a href=
=3D"mailto:[email protected]">[email protected]</a>&gt; wrote:<br></d=
iv><blockquote type=3D"cite" style=3D"margin:0 0 0 .8ex; border-left:2px #7=
29fcf solid;padding-left:1ex">Hello,<br>
<br>
I recently upgrade from Cairo 1.12 to 1.14 (yes, I know these are old<br>
versions), and after doing so noticed a approximately 70% reduction in<br>
performance when rendering scenes that make heavy use of image scaling.<br>
I did some digging and tracking the offending commit down to the<br>
commit: f337342c8 ("V6 image: Use convolution filters for sample<br>
reconstruction when downscaling")<br>
<br>
It appears that this commit is attempting to improve the quality of<br>
downscaled images by implementing new interpolation algorithms in cairo<br>
instead of using the pixman algorithms. My theory is that this is much<br>
slower on ARM processes that have NEON support because pixman has<br>
special implementations of the interpolations algorithms written to<br>
take advantage of NEON, while the new cairo implementations do not.<br>
<br>
Does anyone have any ideas on what a good path forward would be to<br>
restore the ARM+NEON performance? I am planning on trying to reproduce<br>
this with a newer version of cairo to see if it is still a problem, but<br>
I suspect it will be based on the lack of any significant changes in<br>
this code to either cairo or pixman.<br>
<br>
-- <br>
Joshua Watt &lt;<a href=3D"mailto:[email protected]" target=3D"_blank">J=
[email protected]</a>&gt;<br>
<br>
</blockquote></div></blockquote><div><span><pre>-- <br></pre>Joshua Watt &l=
t;<a href=3D"mailto:[email protected]">[email protected]</a>&gt;</spa=
n></div></body></html>

--=-Ugj1gjdK8NLHmuQJAYBx--


--===============1078554132==
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: inline

LS0gCmNhaXJvIG1haWxpbmcgbGlzdApjYWlyb0BjYWlyb2dyYXBoaWNzLm9yZwpodHRwczovL2xp
c3RzLmNhaXJvZ3JhcGhpY3Mub3JnL21haWxtYW4vbGlzdGluZm8vY2Fpcm8=

--===============1078554132==--