Re: MMX version of 4x4 Forward Transform
"Aitor Garay" <[email protected]> Mon, 31 Mar 2003 10:26:26 +0200
| Newsgroups | gmane.comp.video.h264.devel |
|---|---|
| Message-ID | <002101c2f75f$390c33d0$bcb31fac@HAJC0062> |
> I'm trying to write a MMX version of this function. I write it,
> but I do more test before I commit it. It's faster than C++
> version:
That's very interesting. It it helps, each vertical and horizontal row/column is
calculated using the following equations:
T0 = S0 + S3 // Sx: source data
T1 = S1 + S2 // Tx: temporals
T2 = S1 - S2 // Rx: result
T3 = S0 - S3
R0 = T0 + T1
R2 = T0 - T1
R1 = T2 + ( T3 << 1)
R3 = T3 - ( T2 << 1)
Someone commented that the transformation is already quite optimized, and little
gain could be achieved using SIMD. From your timing results, the speed-up is small
( 1.24), but more detailed profiling should be done to assert that.
Do you have experience on SSE/SEE2? Could it be more usefull than MMX?
/AITOR
>
> - C++: 2250ms
> - MMX: 1810ms
>
> This is time showed by openhdot264 when i enable the MMX code. BTW
> This code is not well optimalized... I work on it...
>
> PS. Sorry for my terrible english....
>
> --
> Best regards
> monsti mailto:[email protected]
>
>
>
> -------------------------------------------------------
> This SF.net email is sponsored by:
> The Definitive IT and Networking Event. Be There!
> NetWorld+Interop Las Vegas 2003 -- Register today!
> http://ads.sourceforge.net/cgi-bin/redirect.pl?keyn0001en
> _______________________________________________
> Hdot264-devel mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/hdot264-devel
>
-------------------------------------------------------
This SF.net email is sponsored by: ValueWeb:
Dedicated Hosting for just $79/mo with 500 GB of bandwidth!
No other company gives more support or power for your dedicated server
http://click.atdmt.com/AFF/go/sdnxxaff00300020aff/direct/01/