glReadPixels into PBO is blocking?
"Andras Balogh" <[email protected]> Wed, 07 Feb 2007 22:38:28 -0800
| Newsgroups | gmane.games.devel.opengl |
|---|---|
| Organization | Fat City Network Services, San Diego, California |
| Message-ID | <[email protected]> |
I have to read back framebuffer data (both color and depth) every frame. I don't care about latency, but I don't want to block the CPU at all. For this reason, I've used a ring of PBOs, where I would read into a new PBO every frame, and only touch it after a couple frames later, to avoid blocking. (As discussed in this paper: http://developer.nvidia.com/object/fast_texture_transfers.html) Much to my surprise, this doesn't seem to work efficiently. glReadPixels seems to use up all the CPU, even when reading into a PBO. I have created a bare bones test application: http://www.andrasbalogh.com/gltest.zip to demonstrate it (includes executable, source and project files). The test application renders a single triangle on screen, then reads back a 1x1 pixel area. In this sample application, I don't use the PBO ring, to keep it simple, but I never actually read the data from the PBO. I just read into it, so it should be a 100% asynchronous operation. In the worst case it should decrease framerate, but should not use this much CPU. Commenting out the glReadPixels call, my CPU usage drops to 1%. When I put glFinish() right before calling glReadPixels(), then all the time is spent in glFinish(), and glReadPixels is basically free (I've measured this using RDTSC on a single core computer). This makes me suspect that there's an implicit glFinish() in ReadPixels(), which makes sense, when reading into client memory, but does not, when reading into a PBO. I've also noticed similar behaviour, when I use glCopyTexSubImage2D() to copy into a texture from a FBO. The thing that puzzles me the most, is that if I actually take the glReadPixels() out, and just replace it with a glFinish() call, it still makes my CPU run at a 100%! How is that possible? Even if glFinish() is polling, it should return way before the vsync, and during SwapBuffers() my CPU should be idle. Or am I missing something? I've tested multiple usage types (DYNAMIC_READ, STREAM_READ, STATIC_READ), multiple formats, multiple region sizes, and there's no difference. I've tested on a GeForce6600 and a GeForce7600GS on WinXP with drivers 93.71 and 93.81. Any ideas what I'm doing wrong? Thanks, - Andras ----- FAQ and OpenGL Resources at: http://www.geocities.com/SiliconValley/Hills/9956/OpenGL -- Author: Andras Balogh INET: [email protected] Fat City Hosting, San Diego, California -- http://www.fatcity.com --------------------------------------------------------------------- To REMOVE yourself from this mailing list, send an E-Mail message to: [email protected] (note EXACT spelling of 'ListGuru') and in the message BODY, include a line containing: UNSUB OPENGL-GAMEDEV-L (or the name of mailing list you want to be removed from). You may also send the HELP command for other information (like subscribing).