Re: glReadPixels into PBO is blocking?

"Andras Balogh" <[email protected]> Thu, 08 Feb 2007 22:13:25 -0800
Newsgroups gmane.games.devel.opengl
Organization Fat City Network Services, San Diego, California
Message-ID <[email protected]>
Ok, I've just learned the solution (see discussion here:  
http://www.opengl.org/discussion_boards/ubb/ultimatebb.php?ubb=get_topic;f=3;t=015006):
Apparently, I have to request an 8 bit destination alpha for the  
framebuffer pixelformat, so that it's nice and aligned 32bits per pixel..  
Once I did that, CPU usage dropped to 1%, while executing a glReadPixels  
every frame!

Hats off to nVidia for great dev support!

Thanks,


Andras

On Wed, 07 Feb 2007 23:38:28 -0700, Andras Balogh <[email protected]>  
wrote:

> I have to read back framebuffer data (both color and depth) every frame.  
> I don't care about latency, but I don't want to block the CPU at all.  
> For this reason, I've used a ring of PBOs, where I would read into a new  
> PBO every frame, and only touch it after a couple frames later, to avoid  
> blocking. (As discussed in this paper:  
> http://developer.nvidia.com/object/fast_texture_transfers.html)
>
> Much to my surprise, this doesn't seem to work efficiently. glReadPixels  
> seems to use up all the CPU, even when reading into a PBO. I have  
> created a bare bones test application:  
> http://www.andrasbalogh.com/gltest.zip to demonstrate it (includes  
> executable, source and project files). The test application renders a  
> single triangle on screen, then reads back a 1x1 pixel area. In this  
> sample application, I don't use the PBO ring, to keep it simple, but I  
> never actually read the data from the PBO. I just read into it, so it  
> should be a 100% asynchronous operation. In the worst case it should  
> decrease framerate, but should not use this much CPU.
>
> Commenting out the glReadPixels call, my CPU usage drops to 1%. When I  
> put glFinish() right before calling glReadPixels(), then all the time is  
> spent in glFinish(), and glReadPixels is basically free (I've measured  
> this using RDTSC on a single core computer). This makes me suspect that  
> there's an implicit glFinish() in ReadPixels(), which makes sense, when  
> reading into client memory, but does not, when reading into a PBO.
>
> I've also noticed similar behaviour, when I use glCopyTexSubImage2D() to  
> copy into a texture from a FBO.
>
> The thing that puzzles me the most, is that if I actually take the  
> glReadPixels() out, and just replace it with a glFinish() call, it still  
> makes my CPU run at a 100%! How is that possible? Even if glFinish() is  
> polling, it should return way before the vsync, and during SwapBuffers()  
> my CPU should be idle. Or am I missing something?
>
> I've tested multiple usage types (DYNAMIC_READ, STREAM_READ,  
> STATIC_READ), multiple formats, multiple region sizes, and there's no  
> difference.
> I've tested on a GeForce6600 and a GeForce7600GS on WinXP with drivers  
> 93.71 and 93.81.
>
> Any ideas what I'm doing wrong?
>
> Thanks,
>
>
> - Andras
> -----FAQ and OpenGL Resources at:
>   http://www.geocities.com/SiliconValley/Hills/9956/OpenGL
>


----- 
FAQ and OpenGL Resources at:
  http://www.geocities.com/SiliconValley/Hills/9956/OpenGL

-- 
Author: Andras Balogh
  INET: [email protected]

Fat City Hosting, San Diego, California -- http://www.fatcity.com
---------------------------------------------------------------------
To REMOVE yourself from this mailing list, send an E-Mail message
to: [email protected] (note EXACT spelling of 'ListGuru') and in
the message BODY, include a line containing: UNSUB OPENGL-GAMEDEV-L
(or the name of mailing list you want to be removed from).  You may
also send the HELP command for other information (like subscribing).