Re: multithreading
JVZ <[email protected]>
| Newsgroups | gmane.comp.lib.fox-toolkit.user |
|---|---|
| Message-ID | <[email protected]> |
When you say slow, what sort of runtime are.you looking at? Minutes? Seconds? Milliseconds? On November 20, 2022 6:19:26 PM UTC, John Selverian <[email protected]> wrote: >I should add, the calculation with 1 thread takes about 3 minutes, with >5 threads (# cores- 1) it takes about 4 minutes. > > > > > > > >From: JVZ <[email protected]> >Sent: Sunday, November 20, 2022 12:56 PM >To: [email protected] >Subject: Re: [Foxgui-users] multithreading > > > >Fyi, i just noticed the array will go in and out of scope for each >loop. That might not help with performance... > >On November 20, 2022 4:33:14 PM UTC, Jeroen van der Zijp ><[email protected] <mailto:[email protected]> > wrote: > >On Sun, 20 Nov 2022 10:59:16 -0500 >"John Selverian" <[email protected] ><mailto:[email protected]> > wrote: > >Here is the main part of the code, I left out allocate and >deallocation... > >It is slower then numTreads=12 than for numTreads = 1 > >I have 12 logically processors, the more I use the slower it goes. > > > > > numThreadsRunning = 0; > > > for (FXint i = 0; i < numTreads; i++) > { > FXArray<Material> materialsArray; > >for (int list_box_index = firstListBoxIndex[i]; list_box_index < >lastListBoxIndex[i] + 1; list_box_index++) > { > int j_start = _listBoxDefinition[list_box_index].iStart; > int j_end = _listBoxDefinition[list_box_index].iEnd; > > for (int j = j_start; j < j_end + 1; j++) > { > materialsArray.append(_materialsArray.at(j)); > } > } > > // pass data to thread >checkValuesThreadInitalizedOK[i] = >checkValuesThread[i].passData(materialsArray, >_lastArrayIndexInMatArrayBeforeUserMats, dt); > > if (checkValuesThreadInitalizedOK[i]) > { > checkValuesThread[i].start(); > hThreadCheckValues[i] = checkValuesThread[i].id(); > numThreadsRunning++; > } > } > > > while (true) > { > bool anyThreadsRunning = false; > > for (FXint i = 0; i < numTreads; i++) > { > if (checkValuesThread[i].running()) > { > anyThreadsRunning = true; > break; > } > } > > // if all threads are finished > if (!anyThreadsRunning) > { > break; > } > > // keep UI responsive > if (getApp()->peekEvent()) > { > getApp()->runWhileEvents(); > } > } // while (true) > > > > for (FXint i = 0; i < numTreads; i++) > { > checkValuesThread[i].join(); > } > > >Whether multiple threads is faster or slower depends on a number of >factors: > >1) Is workload big enough to warrant overhead incurred creating >threads, > and does it make up for inter-thread locking etc. needed? > >2) Are workloads localized fighting for common resources that would >make > things slower [e.g. mutexes, memory, etc.]. > >3) Are multiple threads making cache-locality worse? I've had this >exact >scenario, where workloads for threads weren't very large, and the >overhead >of moving "hot" cachelines from main thread to "worker" threads wasn't >worth >the expense. The other cores each have their own L1 and L2, and thus >hot >cachlines from main thread will have to be exchanged to the other >cores. > This causes some cacheline exchanging with other cores. > >4) Why the busy-looping of the main thread? Just use FXMessageChannel >to > let the worker threads know the tasks are done. > >You can also try use FXThreadPool. It can amortize the cost of >thread-startup by >having a number of threads already present, and ready to go when >there's a job. > >With FXThreadPool, threads will not have to be started/stopped for each >parallel >job, but only once at startup. > >Cost of starting a thread, while not huge, is: > > 1) Allocate stack space for the thread, > > 2) Allocate thread-local variables, > > 3) Perform bunch of system calls to launch the thread. > >This cost is negligeable if you perform a multi-hour raytrace or >something >like that. If your entire task is a few seconds, the start-stop cost >could be >signifant. > >For scientific workloads [and I have some experience with that], which >are heavy >matrix number crunching, may have more benefit from vectorization than >parallelization; >I actually would argue to vectorize first, and parallelize second. > >Enable AVX [or AVX512 if you have it], and write simple loops. Use >libraries such >as FFTW3 and OpenBLAS, or Intel Math Kernel Library. These are very >fast and highly >tricked out to max throughput. > >If that still doesn't do it, and you have a GPU, try CUDA [or CUDA >libraries such >as cublas etc.]. > > > > -- JVZ > > > _____ > >Foxgui-users mailing list >[email protected] ><mailto:[email protected]> >https://lists.sourceforge.net/lists/listinfo/foxgui-users > > >-- >Sent from my Android device with K-9 Mail. Please excuse my brevity. -- Sent from my Android device with K-9 Mail. Please excuse my brevity. _______________________________________________ Foxgui-users mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/foxgui-users