[Bug tree-optimization/126372] [17 Regression] 6% slowdown of calculix on Zen2 since r17-2478-g3adb3325954177

"rguenth at gcc dot gnu.org via Gcc-bugs" <[email protected]>
Newsgroups gmane.comp.gcc.bugs
Message-ID <[email protected]/bugzilla/>
https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126372

--- Comment #4 from Richard Biener <rguenth at gcc dot gnu.org> ---
(In reply to Richard Biener from comment #3)
> For -Ofast -g also reproducible on a Zen4 host.  IIRC milc was sensitive to
> code layout changes.  Once I throw perf at it the difference vanishes.
> 
> There is only
> 
> -check_unitarity.c:128:9: optimized: basic block part vectorized using 16
> byte vectors
> -check_unitarity.c:145:9: optimized: basic block part vectorized using 16
> byte vectors
> +check_unitarity.c:146:5: optimized: basic block part vectorized using 16
> byte vectors
> +m_mat_nn.c:88:32: optimized: basic block part vectorized using 16 byte
> vectors
> 
> 
> The m_mat_nn.c (mult_su3_nn) part is us now vectorizing
> 
> void mult_su3_nn( su3_matrix *a, su3_matrix *b, su3_matrix *c ){
>   int i,j;
>   register double t,ar,ai,br,bi,cr,ci;
>     for(i=0;i<3;i++)for(j=0;j<3;j++){
> 
>         ar=a->e[i][0].real; ai=a->e[i][0].imag;
>         br=b->e[0][j].real; bi=b->e[0][j].imag;
>         cr=ar*br; t=ai*bi; cr -= t;
>         ci=ar*bi; t=ai*br; ci += t;
> 
>         ar=a->e[i][1].real; ai=a->e[i][1].imag;
>         br=b->e[1][j].real; bi=b->e[1][j].imag;
>         t=ar*br; cr += t; t=ai*bi; cr -= t;
>         t=ar*bi; ci += t; t=ai*br; ci += t;
> 
>         ar=a->e[i][2].real; ai=a->e[i][2].imag;
>         br=b->e[2][j].real; bi=b->e[2][j].imag;
>         t=ar*br; cr += t; t=ai*bi; cr -= t;  <---  this
>         t=ar*bi; ci += t; t=ai*br; ci += t;  <---
> 
>         c->e[i][j].real=cr;  <--- in addition to this
>         c->e[i][j].imag=ci;
>     }
> }
> 
> vectorizing c->e[1][2].real = <plus reduction>
> 
> +  vectp.627_1209 = &b_10(D)->e[2][2];
> +  vect_bi_120.628_1210 = MEM <vector(2) double> [(double *)vectp.627_1209];
> +  _1213 = BIT_FIELD_REF <vect_bi_120.628_1210, 64, 64>;
> +  _1212 = BIT_FIELD_REF <vect_bi_120.628_1210, 64, 0>;
> ...
> +  vectp.630_1214 = &a_7(D)->e[1][2];
> +  vectp.630_1225 = vectp.630_1214;
> +  vect_ai_323.636_1226 = MEM <vector(2) double> [(double *)vectp.630_1225];
> +  vect_ar_71.637_1227 = VEC_PERM_EXPR <vect_ai_323.636_1226,
> vect_ai_323.636_1226, { 1, 0 }>;
> +  vect_t_611.638_1228 = vect_bi_120.628_1210 * vect_ar_71.637_1227;
> ...
> +  _1229 = .REDUC_PLUS (vect_t_611.638_1228);
> +  _1230 = ci_382 + t_370;
> +  _1231 = _1230 + t_329;
> +  _1232 = _1231 + t_335;
> +  _1233 = _1229 + _1232;
> +  ci_612 = _1233;
> 
> as this has common nodes with the rest it only alters costs slightly,
> increasing profitability even:
> 
> -m_mat_nn.c:90:17: note: Cost model analysis for part in loop 0:
> -  Vector cost: 1804
> -  Scalar cost: 2004
> +  Vector cost: 1928
> +  Scalar cost: 2128
> 
> but then it's probably, overall, not a very profitable vectorization
> of scalar code.
> 
> 
> 
> The check_unitarity.c:check_su3 change looks like a good one, replacing
> three separate plus reductions with a combined MAX reduction and vectorized
> sqrt.

Ignore this ... this was milc.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.