Re: Decoding a web page
Samuel Sieb <[email protected]>
| Newsgroups | gmane.linux.redhat.fedora.general |
|---|---|
| Message-ID | <[email protected]> |
On 6/20/26 9:48 PM, Dave Close wrote: > Every now and then I come across a web page that doesn't seem to follow > conventions. The URL I quote below points to a text file containing > temperatures and other weather data. If I access the URL using curl, > wget, or lynx, I can get the contents of the file but I can't find a > way to get the unrendered content of the page. This matters because > somewhere in that content seems to be the intended name of the file. > I can't seem to find a way to see the name using curl, wget, or lynx. > > But I can see it using Firefox. Accessing that URL with Firefox puts > the file in Firefox's "downloads" directory (however that is defined) > and gives it the proper name. But I want to get the file and its name > with a script. I can't find a way for Firefox to get the file and > then exit without reporting errors and leaving orphaned processes > behind so using Firefox inside a script seems to be impossible. > > Is there any way cleanly to get both the file content and its name > in a bash script? Alternatively, is there a way to see the unrendered > content of the web page? > > Here's the URL (beware line wrapping): > <http://url3286.lacrosseview.com/ls/click?upn=u001.wEwOkX4Mg8d3BzrBweGcjcFmBontLPsLp-2FsIQ1YvgeD6pEs > y-2FqeUzLtpJdkTG4HMKEO0zE1XUZrmEiGYY1H3wDN91LBllaYxAUbrLw5ZmH40hRsXDMJLO-2FMTrTgam7vlhpK-2B1TXlcI3rs > 0TxaUa56Q-3D-3Dx8LQ_xXiUzWINPS0Q7ctV3M9ntHVOnaaHS4EOOecUTAecdk-2Bj4ceVBaLaIxsLbXRRPp0Ne96xQq2y-2B643 > AVFgcu8ZwIJ4BnvUXp0AQZyvRjTqudimAku-2F5abal3l1DdSypjIgVjyrFASld1eEtvMEnVTQPaG5gEGgxQMQYqVIn79pOG3jFQ > RU2JNlB5XVcWIt0pdHOBXeyqSSaNtjl9jF8hUUfc0tF-2BQWQOilBVjd4mIo9qnO0Fm4kZZ0kmQUPGD4gthh20-2FJjthfQ-2FBF > lHmM2QnPmxDA9ZywgzPuqsLi1Nfd3M8NoCTYcy9ejEIrDNS-2Bpv16y7k9fvAu-2BLDz41JlP-2FJoGg-3D-3D> Curl is somehow breaking things. The response to that request is a 302 redirect to another url. You need to add the -L option for curl to follow the redirect. However, if you use curl on that url it gets a 404. I have no idea what it's doing, but it's somehow sending the wrong request. wget works though. wget --content-disposition "long url here" That will download the csv file and give it the right name. If you want that name, you could try grepping the wget output for "Saving". I can't test that now because your token has expired. Probably better to just write the output to a file you specify instead. -- _______________________________________________ users mailing list -- [email protected] To unsubscribe send an email to [email protected] Fedora Code of Conduct: https://docs.fedoraproject.org/en-US/project/code-of-conduct/ List Guidelines: https://fedoraproject.org/wiki/Mailing_list_guidelines List Archives: https://lists.fedoraproject.org/archives/list/[email protected] Do not reply to spam, report it: https://forge.fedoraproject.org/infra/tickets/issues/new