Re: Cannot compile speexdsp 1.2rc3 on ARM64
Frank Barchard <[email protected]> Fri, 29 Jul 2016 17:21:24 -0700
| Newsgroups | gmane.comp.audio.compression.speex.devel |
|---|---|
| Message-ID | <CADdf1xXNbc8pOKivmHf3nSBBzfgAE-H4zbgxLuM1tL0J9waUDw@mail.gmail.com> |
--===============2893175022471004056== Content-Type: multipart/alternative; boundary=001a113e399cedb26b0538cf5851 --001a113e399cedb26b0538cf5851 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: quoted-printable I've filed a bug for aarch64 https://github.com/xiph/speexdsp/issues/7 and provided the port in a fork with a pull request. We need someone to review/merge in the pull request? It provides the source code, but my testing was under Android builds, so there would be some configure changes needed to build it stand alone. On Tue, Apr 19, 2016 at 4:32 PM, Frank Barchard <[email protected]> wrote: > Hi I'm new to speex list but joined because I'm needing to port the Neo= n > to ARM64. > On that function, saturate_32bit_to_16bit(), I noticed the ifdef's are > wrong. > The first version is for normal arm 32 bit arm and should be used for > arm32 and thumb2 but not thumb1. > The second version is 32 bit neon and should be #ifdef __ARM_NEON__ > I've done a third version which is 64 bit neon. I'm working off an > android version which is rc2 so I'll need to integrate, but here it is: > > #if defined(__aarch64__) > static inline int32_t saturate_32bit_to_16bit(int32_t a) { > int32_t ret; > asm volatile ("sqxtn h0, %s[a]\n" > "sxtl v0.4s, v0.4h\n" > "fmov %w[ret], s0\n" > : [ret] "=3D&r" (ret) > : [a] "w" (a) > : "v0" ); > return ret; > } > #elif defined(__ARM_NEON__) > static inline int32_t saturate_32bit_to_16bit(int32_t a) { > int32_t ret; > asm volatile ("vmov.s32 d24[0], %[a]\n" > "vqmovn.s32 d24, q12\n" > "vmov.s16 %[ret], d24[0]\n" > : [ret] "=3D&r" (ret) > : [a] "r" (a) > : "q12", "d24", "d25" ); > return ret; > } > #else > static inline int32_t saturate_32bit_to_16bit(int32_t a) { > return max(-32768, min(32767, a)); > } > #endif > > To test it I wrote a stand alone test and ran it via adb. > Anyone able to help with review/integration? > There are 4 functions in resample_neon.h thats just the first/easiest. > > > > On Sat, Mar 28, 2015 at 11:28 AM, Evan JIANG <[email protected]> wrote= : > >> Hi all, >> I build successfully with speex-1.2rc2. And with speexdsp 1.2rc3, = I >> build with i386, X86_64, armv7 and armv7s all passed. >> But when I build for ARM64 (for iPhone 6), it failed with: >> /Applications/Xcode.app/Contents/Developer/usr/bin/make all-recursive >> Making all in libspeexdsp >> CC preprocess.lo >> CC jitter.lo >> CC mdf.lo >> CC fftwrap.lo >> CC filterbank.lo >> CC resample.lo >> In file included from resample.c:104: >> ./resample_neon.h:134:12: error: unknown register name 'q0' in asm >> : "q0"); >> ^ >> ./resample_neon.h:195:13: error: invalid output constraint '+l' in asm >> [len] "+l" (len), [remainder] "+l" (remainder) >> ^ >> 2 errors generated. >> make[2]: *** [resample.lo] Error 1 >> make[1]: *** [all-recursive] Error 1 >> make: *** [all] Error 2 >> >> >> As I googled out, I found it's said: >> >> arm64 has a totally different instruction set. >> >> See: http://people.linaro.org/~rikuvoipio/aarch64-talk/ >> >> The NEON assembly code needs a rewrite. >> >> >> >> But I'm not familiar with ASM code. Could anyone help to fix that? >> >> Best regards, >> Evan JIANG >> >> _______________________________________________ >> Speex-dev mailing list >> [email protected] >> http://lists.xiph.org/mailman/listinfo/speex-dev >> >> > --001a113e399cedb26b0538cf5851 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr">I've filed a bug for aarch64<div><a href=3D"https://= github.com/xiph/speexdsp/issues/7">https://github.com/xiph/speexdsp/issue= s/7</a><br></div><div><br></div><div>and provided the port in a fork with= a pull request.=C2=A0 We need someone to review/merge in the pull reques= t?</div><div>It provides the source code, but my testing was under Androi= d builds, so there would be some configure changes needed to build it sta= nd alone.</div></div><div class=3D"gmail_extra"><br><div class=3D"gmail_q= uote">On Tue, Apr 19, 2016 at 4:32 PM, Frank Barchard <span dir=3D"ltr">&= lt;<a href=3D"mailto:[email protected]" target=3D"_blank">fbarchard@go= ogle.com</a>></span> wrote:<br><blockquote class=3D"gmail_quote" style= =3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><div d= ir=3D"ltr">Hi I'm new to speex list but joined because I'm needin= g to port the Neon to ARM64.<div>On that function, saturate_32bit_to_16bi= t(),=C2=A0I noticed the <!-- -->ifdef's are wrong.</div><div>The first version is for normal arm 3= 2 bit arm and should be used for arm32 and thumb2 but not thumb1.</div><d= iv>The second version is 32 bit neon and should be #ifdef __ARM_NEON__</d= iv><div>I've done a third version which is 64 bit neon. =C2=A0 I'= m working off an android version which is rc2 so I'll need to integra= te, but here it is:</div><div><br></div><div><div>#if defined(__aarch64__= )</div><div>static inline int32_t saturate_32bit_to_16bit(int32_<wbr>t a)= {</div><div><span class=3D"m_-3490218015663917727gmail-Apple-tab-span" s= tyle=3D"white-space:pre-wrap"> </span>int32_t ret;</div><div><span class=3D= "m_-3490218015663917727gmail-Apple-tab-span" style=3D"white-space:pre-wra= p"> </span>asm volatile ("sqxtn h0, %s[a]\n"</div><div><span cl= ass=3D"m_-3490218015663917727gmail-Apple-tab-span" style=3D"white-space:p= re-wrap"> </span> =C2=A0 =C2=A0 =C2=A0"sxtl =C2=A0v0.4s, v0.4h\n&qu= ot;</div><div><!-- --><span class=3D"m_-3490218015663917727gmail-Apple-tab-span" style=3D"wh= ite-space:pre-wrap"> </span> =C2=A0 =C2=A0 =C2=A0"fmov %w[ret], s0\= n"</div><div><span class=3D"m_-3490218015663917727gmail-Apple-tab-sp= an" style=3D"white-space:pre-wrap"> </span> =C2=A0 =C2=A0 =C2=A0: [ret] = "=3D&r" (ret)</div><div><span class=3D"m_-34902180156639177= 27gmail-Apple-tab-span" style=3D"white-space:pre-wrap"> </span> =C2=A0 =C2= =A0 =C2=A0: [a] "w" (a)</div><div><span class=3D"m_-34902180156= 63917727gmail-Apple-tab-span" style=3D"white-space:pre-wrap"> </span> =C2= =A0 =C2=A0 =C2=A0: "v0" );</div><div><span class=3D"m_-34902180= 15663917727gmail-Apple-tab-span" style=3D"white-space:pre-wrap"> </span>r= eturn ret;</div><div>}</div><div>#elif defined(__ARM_NEON__)</div><div>st= atic inline int32_t saturate_32bit_to_16bit(int32_<wbr>t a) {</div><div><= span class=3D"m_-3490218015663917727gmail-Apple-tab-span" style=3D"white-= space:pre-wrap"> </span>int32_t ret;</div><div><!-- --><span class=3D"m_-3490218015663917727gmail-Apple-tab-span" style=3D"wh= ite-space:pre-wrap"> </span>asm volatile ("vmov.s32 d24[0], %[a]\n&q= uot;</div><div><span class=3D"m_-3490218015663917727gmail-Apple-tab-span"= style=3D"white-space:pre-wrap"> </span> =C2=A0 =C2=A0 =C2=A0"vqmov= n.s32 d24, q12\n"</div><div><span class=3D"m_-3490218015663917727gma= il-Apple-tab-span" style=3D"white-space:pre-wrap"> </span> =C2=A0 =C2=A0= =C2=A0"vmov.s16 %[ret], d24[0]\n"</div><div><span class=3D"m_-= 3490218015663917727gmail-Apple-tab-span" style=3D"white-space:pre-wrap"> = </span> =C2=A0 =C2=A0 =C2=A0: [ret] "=3D&r" (ret)</div><di= v><span class=3D"m_-3490218015663917727gmail-Apple-tab-span" style=3D"whi= te-space:pre-wrap"> </span> =C2=A0 =C2=A0 =C2=A0: [a] "r" (a)<= /div><div><span class=3D"m_-3490218015663917727gmail-Apple-tab-span" styl= e=3D"white-space:pre-wrap"> </span> =C2=A0 =C2=A0 =C2=A0: "q12"= ;, "d24", "d25" );</div><div><!-- --><span class=3D"m_-3490218015663917727gmail-Apple-tab-span" style=3D"wh= ite-space:pre-wrap"> </span>return ret;</div><div>}</div><div>#else</div>= <div>static inline int32_t saturate_32bit_to_16bit(int32_<wbr>t a) {</div= ><div><span class=3D"m_-3490218015663917727gmail-Apple-tab-span" style=3D= "white-space:pre-wrap"> </span>return max(-32768, min(32767, a));</div><d= iv>}</div><div>#endif</div></div><div><br></div><div>To test it I wrote a= stand alone test and ran it via adb.</div><div>Anyone able to help with = review/integration?</div><div>There are 4 functions in resample_neon.h th= ats just the first/easiest.</div><div><br></div><div><br></div></div><div= class=3D"gmail_extra"><br><div class=3D"gmail_quote"><div><div class=3D"= h5">On Sat, Mar 28, 2015 at 11:28 AM, Evan JIANG <span dir=3D"ltr"><<a= href=3D"mailto:[email protected]" target=3D"_blank">[email protected]<= /a>></span> wrote:<br></div></div><!-- --><blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-le= ft:1px #ccc solid;padding-left:1ex"><div><div class=3D"h5"><div dir=3D"lt= r"><div><div><div><div><div><div>Hi all,<br></div>=C2=A0=C2=A0=C2=A0 I bu= ild successfully with speex-1.2rc2. And with speexdsp 1.2rc3, I build wit= h i386, X86_64, armv7 and armv7s all passed.<br></div>=C2=A0=C2=A0 But wh= en I build for ARM64 (for iPhone 6), it failed with:<br>/Applications/Xco= de.app/Conten<wbr>ts/Developer/usr/bin/make=C2=A0 all-recursive<br>Making= all in libspeexdsp<br>=C2=A0 CC=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 prep= rocess.lo<br>=C2=A0 CC=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 jitter.lo<br>=C2= =A0 CC=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 mdf.lo<br>=C2=A0 CC=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0 fftwrap.lo<br>=C2=A0 CC=C2=A0=C2=A0=C2=A0=C2=A0=C2= =A0=C2=A0 filterbank.lo<br>=C2=A0 CC=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 = resample.lo<br>In file included from resample.c:104:<br>./resample_neon.h= :134:12: error: unknown register name 'q0' in asm<br>=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 : "q0");<br>=C2=A0=C2=A0=C2= =A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 ^<br>./resample_neon.h:195:= 13: error: invalid output constraint '+l<!-- -->' in asm<br>=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 [len] "= +l" (len), [remainder] "+l" (remainder)<br>=C2=A0=C2=A0=C2= =A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 ^<br>2 error= s generated.<br>make[2]: *** [resample.lo] Error 1<br>make[1]: *** [all-r= ecursive] Error 1<br>make: *** [all] Error 2<br><br><br></div>As I google= d out, I found it's said:<br><br><pre>arm64 has a totally different i= nstruction set. See: <a href=3D"http://people.linaro.org/%7Erikuvoipio/aarch64-talk/" rel= =3D"nofollow" target=3D"_blank">http://people.linaro.org/~riku<wbr>voipio= /aarch64-talk/</a> The NEON assembly code needs a rewrite.</pre><br><br></div>But I'm no= t familiar with ASM code. Could anyone help to fix that?<br><br></div>Bes= t regards,<br></div>Evan JIANG<br></div> <br></div></div>______________________________<wbr>_________________<br> Speex-dev mailing list<br> <a href=3D"mailto:[email protected]" target=3D"_blank">[email protected]= g</a><br> <a href=3D"http://lists.xiph.org/mailman/listinfo/speex-dev" rel=3D"noref= errer" target=3D"_blank">http://lists.xiph.org/mailman/<wbr>listinfo/spee= x-dev</a><br> <br></blockquote></div><br></div> </blockquote></div><br></div> --001a113e399cedb26b0538cf5851-- --===============2893175022471004056== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KU3BlZXgtZGV2 IG1haWxpbmcgbGlzdApTcGVleC1kZXZAeGlwaC5vcmcKaHR0cDovL2xpc3RzLnhpcGgub3JnL21h aWxtYW4vbGlzdGluZm8vc3BlZXgtZGV2Cg== --===============2893175022471004056==--