Re: NullPointerException in Dfs.resolve
"Vella, Shon" <[email protected]> Mon, 15 Jun 2015 08:46:51 -0600
| Newsgroups | gmane.network.samba.java |
|---|---|
| Message-ID | <CAND51tQSUJEvhXkX8ZEY88FiGgEXPvndasAx1YGwRGzRODDV5A@mail.gmail.com> |
--20cf3079b68832a19e05188f8705 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: quoted-printable One of the patches from google submitted to the list last year: http://article.gmane.org/gmane.network.samba.java/9410 has an attempted fix of this. https://code.google.com/p/google-enterprise-connector-file-system/source/de= tail?r=3D563 It seems to work for me - in any case it shouldn't ever throw that NPE out of resolveDfs. *Shon Vella* *Identity Automation* Product Engineer 281-220-0021 x2030 office 281-817-5579 fax www.identityautomation.com On Mon, Jun 15, 2015 at 5:45 AM, Martin Kutter <[email protected]> wrote: > Thanks for your reply. > I've tested 1.3.18b - it does not fix this specific error. > > Regarding the fix suggested by Conrad: In the meantime, I've also > implemented a fix for the issue - I changed SmbFile.resolveDfs() > (lines 671 and following) to > > DfsReferral dr =3D null; > // disconnect is synchronized to transport, too. > // make sure our transport doesn't get disconnected while > // we're inside the synchronized block > synchronized (tree.session.transport) { > if (tree.session.transport.tconHostName =3D=3D null) { > // disconnect properly if connection is lost > tree.treeDisconnect(false); > } > tree.session.transport.connect(); > dr =3D dfs.resolve(tree.session.transport.tconHostName, > tree.share, > unc, > auth); > } > if (dr !=3D null) { > > > As my knowledge of CIFS is quite limited, I have no idea whether this > is right (or has bad side effects). > > The idea behind is similar (but not equal) to Conrad's fix: In case the > transport has disconnected, disconnect properly and reconnect. > > Best regards, > > Martin > > On Sun, 14 Jun 2015 11:37:48 -0400, Michael B Allen <[email protected]> > wrote: > > I have added this post to the list of people who have reported it to > > the TODO list so that it can be considered when I get around to > > looking at this. > > > > Note that the 1.3.18b mentioned in the link cited is here: > > > > http://jcifs.samba.org/old/jcifs-1.3.18b.jar > > > > Although I cannot recall what it actually does anymore it might be > > worth a try. We never received feedback about it. > > > > Mike > > > > On Sun, Jun 14, 2015 at 5:00 AM, Conrad Herrmann <[email protected]> > > wrote: > >> Martin, > >> > >> > >> > >> I have recently run into the same problem. > >> > >> > >> > >> I think the problem is that SmbFile.resolveDfs() uses the currently > >> connected transport as the DFS resolver/domain server > >> (tree.session.transport.tconHostName), but it is very possible that > >> there is > >> no currently connected transport. In your case, that happens when the > >> file > >> server forces the TCP connection to close, and the transport tears > itself > >> down. > >> > >> > >> > >> And, although the resolveDfs() method calls connect0(), in fact this > does > >> nothing because doConnect() doesn't force creation of a new connection > >> if we > >> are talking about a DFS resolved path. > >> > >> > >> > >> It seems to me that in that case, we have to start over again at the > top > >> of > >> the referral tree, with the Domain. > >> > >> > >> > >> So my solution has this: change the code for SmbFile.resolveDfs() > lines > >> 671 > >> (or so) so that it says: > >> > >> > >> > >> connect0(); > >> > >> > >> > >>> String hostName =3D tree.session.transport.tconHostName; > >> > >>> String domainDfsServerName =3D getServerWithDfs(); > >> > >>> if (hostName =3D=3D null) > >> > >>> hostName =3D domainDfsServerName; > >> > >> > >> > >> DfsReferral dr =3D dfs.resolve( > >> > >>> hostName, > >> > >> tree.share, > >> > >> unc, > >> > >> auth); > >> > >> > >> > >> The code comes from the other use of tconHostName, in > >> SmbFile.doConnect(): > >> > >> String hostName =3D getServerWithDfs(); > >> > >> tree.inDomainDfs =3D dfs.resolve(hostName, tree.share, null, > auth) > >> !=3D > >> null; > >> > >> In this code, we are getting the DFS resolver (which might be the > domain > >> server) as the hostName, and asking it to resolve our share. > >> > >> > >> > >> Basically what this new code is saying is that: > >> > >> - in the case where the transport has closed (ie, because of a timeout > or > >> TCP close on the DFS server side) reconnect to the DFS domain server i= n > >> order to resolve a share's DFS server. > >> > >> > >> > >> I can imagine a case where this doesn't work--if we have multiple > levels > >> of > >> DFS redirection, where the domain server cannot redirect the client to > a > >> deep subdirectory. But, I don't even know if this is possible in DFS. > >> If > >> it is, then at least this solution removes the top level case and > >> identifies > >> the problem, which would require walking down the DFS resolution path > to > >> resolve the actual file server. > >> > >> > >> > >> Conrad Herrmann > >> > >> Primdaesk, Inc. > >> > >> > >> > >>> Hi, > >> > >>> > >> > >>> I encountered a NullpointerException similar to > >> > >>> https://lists.samba.org/archive/jcifs/2012-January/009856.html - at > >> > >>> least the stack traces are similar. > >> > >>> > >> > >>> My Environment (Client side): > >> > >>> - jcifs 1.3.18 > >> > >>> - IBM JDK 7 > >> > >>> - AIX 7.1 > >> > >>> > >> > >>> The NPE occured in a (in-house) plugin for the Jenkins build server. > In > >> > >>> this system, JCIFs is used to recursively copy files from a Windows > >> > >>> share to an AIX machine. > >> > >>> > >> > >>> Re-running a build shortly after it finished triggered the NPE. > >> > >>> > >> > >>> After some debugging, it seems to me like the SmbFile=E2=80=99s under= lying > >> > >>> transport is closed (by timeout), and when SmbFile.resolveDfs is > called, > >> > >>> the transport is not reconnected (unlike, for example, later in > >> > >>> SmbFile.resolve, or in SmbSession.getChallenge). > >> > >>> > >> > >>> I was able to reproduce the NPE during debugging using the following > >> > >>> steps: > >> > >>> - Trigger a build (recursively copying from a CIFS DFS tree) > >> > >>> - Wait until the transport objects disconnect by timeout (tracked by > >> > >>> breakpoint) > >> > >>> - Retrigger the build (recursively copying the same directory > structure) > >> > >>> > >> > >>> The Jenkins plugin usually runs the second JCIFS copy operation in th= e > >> > >>> same thread than the first (though that's not guaranteed). > >> > >>> Each run uses a new SmbFile object. > >> > >>> > >> > >>> Am I missing something (like some close operation on SmbFile)? > >> > >>> Is this a known error? > >> > >>> Can I do something to fix it? > >> > >>> > >> > >>> Best regards, > >> > >>> > >> > >>> Martin > >> > >>> > --20cf3079b68832a19e05188f8705 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr">One of the patches from google submitted to the list last = year:<div><br></div><div><a href=3D"http://article.gmane.org/gmane.network.= samba.java/9410">http://article.gmane.org/gmane.network.samba.java/9410</a>= <br></div><div><br></div><div>has an attempted fix of this.</div><div><br><= /div><div><a href=3D"https://code.google.com/p/google-enterprise-connector-= file-system/source/detail?r=3D563">https://code.google.com/p/google-enterpr= ise-connector-file-system/source/detail?r=3D563</a><br></div><div><br></div= ><div>It seems to work for me - in any case it shouldn't ever throw tha= t NPE out of resolveDfs.</div><div><br></div></div><div class=3D"gmail_extr= a"><br clear=3D"all"><div><div class=3D"gmail_signature"><div dir=3D"ltr"><= div style=3D"font-size:15px;padding-left:5px"><strong><font color=3D"#00000= 0">Shon Vella</font></strong><br></div><div style=3D"font-size:15px;color:r= gb(229,37,38);padding-left:5px"><b>Identity Automation</b></div><div style= =3D"font-size:15px;padding-left:5px"><font color=3D"#000000">Product Engine= er</font></div><div style=3D"font-size:15px;padding-left:5px">281-220-0021 = x2030=C2=A0<span style=3D"color:rgb(229,37,38)">office</span></div><div sty= le=3D"font-size:15px;padding-left:5px"><a href=3D"tel:281-817-5579" value= =3D"+12818175579" style=3D"color:rgb(17,85,204)" target=3D"_blank">281-817-= 5579</a>=C2=A0<span style=3D"color:rgb(229,37,38)">fax</span></div><div sty= le=3D"font-size:15px;padding-left:5px"><a href=3D"http://www.identityautoma= tion.com/" style=3D"color:rgb(17,85,204)" target=3D"_blank">www.identityaut= omation.com</a></div></div></div></div> <br><div class=3D"gmail_quote">On Mon, Jun 15, 2015 at 5:45 AM, Martin Kutt= er <span dir=3D"ltr"><<a href=3D"mailto:[email protected]" target= =3D"_blank">[email protected]</a>></span> wrote:<br><blockquote c= lass=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;= padding-left:1ex">Thanks for your reply.<br> I've tested 1.3.18b - it does not fix this specific error.<br> <br> Regarding the fix suggested by Conrad: In the meantime, I've also<br> implemented a fix for the issue - I changed SmbFile.resolveDfs()<br> (lines 671 and following) to<br> <br> DfsReferral dr =3D null;<br> // disconnect is synchronized to transport, too.<br> // make sure our transport doesn't get disconnected while<br> // we're inside the synchronized block<br> synchronized (tree.session.transport) {<br> =C2=A0 =C2=A0 if (tree.session.transport.tconHostName =3D=3D null) {<br> =C2=A0 =C2=A0 =C2=A0 =C2=A0 // disconnect properly if connection is lost<br= > =C2=A0 =C2=A0 =C2=A0 =C2=A0 tree.treeDisconnect(false);<br> =C2=A0 =C2=A0 }<br> =C2=A0 =C2=A0 tree.session.transport.connect();<br> =C2=A0 =C2=A0 dr =3D dfs.resolve(tree.session.transport.tconHostName,<br> =C2=A0 =C2=A0 =C2=A0 =C2=A0 tree.share,<br> =C2=A0 =C2=A0 =C2=A0 =C2=A0 unc,<br> =C2=A0 =C2=A0 =C2=A0 =C2=A0 auth);<br> }<br> if (dr !=3D null) {<br> <br> <br> As my knowledge of CIFS is quite limited, I have no idea whether this<br> is right (or has bad side effects).<br> <br> The idea behind is similar (but not equal) to Conrad's fix: In case the= <br> transport has disconnected, disconnect properly and reconnect.<br> <br> Best regards,<br> <br> Martin<br> <br> On Sun, 14 Jun 2015 11:37:48 -0400, Michael B Allen <<a href=3D"mailto:i= [email protected]">[email protected]</a>><br> wrote:<br> <div class=3D"HOEnZb"><div class=3D"h5">> I have added this post to the = list of people who have reported it to<br> > the TODO list so that it can be considered when I get around to<br> > looking at this.<br> ><br> > Note that the 1.3.18b mentioned in the link cited is here:<br> ><br> >=C2=A0 =C2=A0 <a href=3D"http://jcifs.samba.org/old/jcifs-1.3.18b.jar" = rel=3D"noreferrer" target=3D"_blank">http://jcifs.samba.org/old/jcifs-1.3.1= 8b.jar</a><br> ><br> > Although I cannot recall what it actually does anymore it might be<br> > worth a try. We never received feedback about it.<br> ><br> > Mike<br> ><br> > On Sun, Jun 14, 2015 at 5:00 AM, Conrad Herrmann <<a href=3D"mailto= :[email protected]">[email protected]</a>><br> > wrote:<br> >> Martin,<br> >><br> >><br> >><br> >> I have recently run into the same problem.<br> >><br> >><br> >><br> >> I think the problem is that SmbFile.resolveDfs() uses the currentl= y<br> >> connected transport as the DFS resolver/domain server<br> >> (tree.session.transport.tconHostName), but it is very possible tha= t<br> >> there is<br> >> no currently connected transport.=C2=A0 In your case, that happens= when the<br> >> file<br> >> server forces the TCP connection to close, and the transport tears= <br> itself<br> >> down.<br> >><br> >><br> >><br> >> And, although the resolveDfs() method calls connect0(), in fact th= is<br> does<br> >> nothing because doConnect() doesn't force creation of a new co= nnection<br> >> if we<br> >> are talking about a DFS resolved path.<br> >><br> >><br> >><br> >> It seems to me that in that case, we have to start over again at t= he<br> top<br> >> of<br> >> the referral tree, with the Domain.<br> >><br> >><br> >><br> >> So my solution has this:=C2=A0 change the code for SmbFile.resolve= Dfs()<br> lines<br> >> 671<br> >> (or so) so that it says:<br> >><br> >><br> >><br> >>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0connect0();<br> >><br> >><br> >><br> >>>=C2=A0 =C2=A0 =C2=A0 =C2=A0 String hostName =3D tree.session.tr= ansport.tconHostName;<br> >><br> >>>=C2=A0 =C2=A0 =C2=A0 =C2=A0 String domainDfsServerName =3D getS= erverWithDfs();<br> >><br> >>>=C2=A0 =C2=A0 =C2=A0 =C2=A0 if (hostName =3D=3D null)<br> >><br> >>>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 hostName =3D domainDf= sServerName;<br> >><br> >><br> >><br> >>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0DfsReferral dr =3D dfs.resolve(<b= r> >><br> >>>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 = =C2=A0 hostName,<br> >><br> >>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2= =A0 =C2=A0tree.share,<br> >><br> >>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2= =A0 =C2=A0unc,<br> >><br> >>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2= =A0 =C2=A0auth);<br> >><br> >><br> >><br> >> The code comes from the other use of tconHostName, in<br> >> SmbFile.doConnect():<br> >><br> >>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0String hostName =3D getServerWith= Dfs();<br> >><br> >>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0tree.inDomainDfs =3D dfs.resolve(= hostName, tree.share, null,<br> auth)<br> >>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0!=3D<br> >> null;<br> >><br> >> In this code, we are getting the DFS resolver (which might be the<= br> domain<br> >> server) as the hostName, and asking it to resolve our share.<br> >><br> >><br> >><br> >> Basically what this new code is saying is that:<br> >><br> >> - in the case where the transport has closed (ie, because of a tim= eout<br> or<br> >> TCP close on the DFS server side) reconnect to the DFS domain serv= er in<br> >> order to resolve a share's DFS server.<br> >><br> >><br> >><br> >> I can imagine a case where this doesn't work--if we have multi= ple<br> levels<br> >> of<br> >> DFS redirection, where the domain server cannot redirect the clien= t to<br> a<br> >> deep subdirectory.=C2=A0 But, I don't even know if this is pos= sible in DFS.<br> >> If<br> >> it is, then at least this solution removes the top level case and<= br> >> identifies<br> >> the problem, which would require walking down the DFS resolution p= ath<br> to<br> >> resolve the actual file server.<br> >><br> >><br> >><br> >> Conrad Herrmann<br> >><br> >> Primdaesk, Inc.<br> >><br> >><br> >><br> >>> Hi,<br> >><br> >>><br> >><br> >>> I encountered a NullpointerException similar to<br> >><br> >>> <a href=3D"https://lists.samba.org/archive/jcifs/2012-January/= 009856.html" rel=3D"noreferrer" target=3D"_blank">https://lists.samba.org/a= rchive/jcifs/2012-January/009856.html</a> - at<br> >><br> >>> least the stack traces are similar.<br> >><br> >>><br> >><br> >>> My Environment (Client side):<br> >><br> >>> - jcifs 1.3.18<br> >><br> >>> - IBM JDK 7<br> >><br> >>> - AIX 7.1<br> >><br> >>><br> >><br> >>> The NPE occured in a (in-house) plugin for the Jenkins build s= erver.<br> In<br> >><br> >>> this system, JCIFs is used to recursively copy files from a Wi= ndows<br> >><br> >>> share to an AIX machine.<br> >><br> >>><br> >><br> >>> Re-running a build shortly after it finished triggered the NPE= .<br> >><br> >>><br> >><br> >>> After some debugging, it seems to me like the SmbFile=E2=80=99= s underlying<br> >><br> >>> transport is closed (by timeout), and when SmbFile.resolveDfs = is<br> called,<br> >><br> >>> the transport is not reconnected (unlike, for example, later i= n<br> >><br> >>> SmbFile.resolve, or in SmbSession.getChallenge).<br> >><br> >>><br> >><br> >>> I was able to reproduce the NPE during debugging using the fol= lowing<br> >><br> >>> steps:<br> >><br> >>> - Trigger a build (recursively copying from a CIFS DFS tree)<b= r> >><br> >>> - Wait until the transport objects disconnect by timeout (trac= ked by<br> >><br> >>> breakpoint)<br> >><br> >>> - Retrigger the build (recursively copying the same directory<= br> structure)<br> >><br> >>><br> >><br> >>> The Jenkins plugin usually runs the second JCIFS copy operatio= n in the<br> >><br> >>> same thread than the first (though that's not guaranteed).= <br> >><br> >>> Each run uses a new SmbFile object.<br> >><br> >>><br> >><br> >>> Am I missing something (like some close operation on SmbFile)?= <br> >><br> >>> Is this a known error?<br> >><br> >>> Can I do something to fix it?<br> >><br> >>><br> >><br> >>> Best regards,<br> >><br> >>><br> >><br> >>> Martin<br> >><br> >>><br> </div></div></blockquote></div><br></div> --20cf3079b68832a19e05188f8705--