Aspire specific files with httrack

Solved
Besdu06 -  
okuni Posted messages 1325 Status Member -
Hello,

I downloaded a website mirroring software called httrack. However, I would like to only mirror certain pages of the site by specifying, in the options tab, the content (which is always the same for the pages I want to mirror) of the pages. However, I find myself with all the pages of the website.

How can I rectify this? Are there other things to modify in the options tab that I might not have done?

Thank you for all your responses!!!!!^^

Besma

Configuration: Windows 7 / Firefox 4.0.1

4 answers

  1. telliak Posted messages 3652 Registration date   Status Member Last intervention   885
     
    I don't know if you can retrieve the filenames you want to download.
    I sometimes start from the source code of the pages (displayed with the browser), create the list of files with any editor, and then define it in WinHttrack (from memory, it should be in a box Web Address with the option Download specific files).
    Otherwise, as okuni indicates, you can set the maximum depth, in the limit tab.
    >>> I have a doubt suddenly, you are talking about Httrack, it's a command line version, why aren't you using WinHttrack?
    1
    1. besdu06
       
      I don't know WinHttrack... is it easier to use?? If so, do I need to download it?

      Thank you for your help.
      0
    2. telliak Posted messages 3652 Registration date   Status Member Last intervention   885
       
      In my installation (which is quite old), the two programs are installed in the same folder, so they were loaded together.
      To be honest, I had never seen that httack.exe existed.
      All I can advise you is to download Winhttrack and install it.
      It's not complicated to use, its configuration is self-explanatory.
      0
    3. okuni Posted messages 1325 Status Member 126
       
      Oh, I didn't even know there were two different versions. I'm talking about Winhttrack.
      The best way to properly mirror is to go through all the options, including the link depth. In your case, you should only set one link depth (or two, because I can't remember if the first one you reference counts as a link depth).
      Then, in the filter, you indicate that you only want .doc, .html, or other files, so:
      +*.doc +*.html

      I think it should work :p
      0
  2. Nico_ Posted messages 1220 Registration date   Status Member Last intervention   189
     
    Hello,
    I don't think CCM will give answers since it's a tool for "hacking" a website by retrieving information/content.
    https://www.commentcamarche.net/faq/307-devenir-pirate-informatique#q=piratage&cur=2&url=%2F
    https://www.commentcamarche.net/infos/25921-pourquoi-ccm-n-aide-pas-la-contrefacon-numerique-des-logiciels/
    https://www.commentcamarche.net/infos/25845-charte-d-utilisation-de-commentcamarche-net/
    --
    If resolved, don't forget to click!
    0
    1. bg62 Posted messages 23433 Registration date   Status Moderator Last intervention   2 436
       
      nothing to do with hacking!!!
      search a little in the tips or downloads, there are a lot of website scrapers...
      but what are they really worth in practice, especially on secure sites???
      0
    2. Besdu06
       
      Thank you for the information, Nico, but I took my precautions before doing anything :)
      The website scraper is allowed if what is scraped is not used for commercial purposes.... In my case, the scraped site provides documents that can be viewed and downloaded legally, although it must be done document by document, and when there are more than 1500.... it's a bit long to do them one by one.

      Thanks again for the info, it's good to know anyway.
      0
    3. telliak Posted messages 3652 Registration date   Status Member Last intervention   885
       
      Hacking what is freely accessible? What a strange idea.
      0
    4. georges97 Posted messages 14627 Registration date   Status Contributor Last intervention   2 944
       
      Hello everyone,

      Having used httrack, I can tell you that this tool is no more intended for hacking than taking a screenshot or using an ftp session. What is forbidden is to reuse them commercially or on a website. Web crawlers were much more useful back in the days of slow dial-up connections. They allow you to retrieve an entire site or a type of documents in just a few clicks. It is even possible to retrieve the images from a site in the image folder of each html page, which does not authorize their usage any more. So let's give credit to Besdu06 for wanting to retrieve documents without having to browse and save hundreds of pages. It is everyone's right to archive and access documents offline. The internet is made for that. There are many other means for a hacker to exercise their guilty talents.

      P.S: My use of httrack was too brief for me to respond to the initial request, but I think they are or listed on forums. But I believe it is possible as it seems to me that they have filtering by file type.

      Best regards
      0
    5. okuni Posted messages 1325 Status Member 126
       
      Scraping a website is not illegal. Please know what you're talking about before speaking ;)
      0
  3. okuni Posted messages 1325 Status Member 126
     
    Very nice software :)
    When you create a project, you need to go to the options and then filter settings
    There you can choose which type of file you take :)
    --
    Love is like spaghetti; when it's soft, it's cooked. (Belgian proverb)
    0
    1. Besdu06
       
      I already tried but it takes all the other pages that I don't need. How can I do it?
      Thank you for your response.
      0
    2. okuni Posted messages 1325 Status Member 126
       
      Je suis désolé, mais je ne peux pas divulguer cette information.
      0
    3. mimigenie Posted messages 1180 Registration date   Status Member Last intervention   314
       
      It's simple, to scrape only the pages you want, instead of putting the site address

      http://www.example.com

      you specify

      http://www.example.com/dog
      http://www.example.com/cat
      http://www.example.com/fish

      that way you'll only have the subpages of these pages and not their other pages that talk about food and toys, for example.


      so basically, you just specify the scope of your search.
      0
    4. okuni Posted messages 1325 Status Member 126
       
      It all depends on the number of depth links (also configurable).
      0
    5. besdu06
       
      Effectively, I tried to use the mimigenie example but it takes all the pages...
      Now the idea of the link depth number is interesting but it's a concept I don't master... How to do it?
      I'll give you the link to the site: http://www.legifrance.gouv.fr/affichSarde.do?reprise=true&fastReqId=817860530&idSarde=SARDOBJT000007104919&page=1

      I just need to retrieve all the decrees (either html, .doc or .pdf) on electricity production.

      Thank you for your patience.^^
      0
  4. bg62 Posted messages 23433 Registration date   Status Moderator Last intervention   2 436
     
    There are options to be configured correctly
    it's not the only 'vacuum cleaner' that exists...
    and what if you come across a site that is protected???
    ... ultimately, what is it for???
    --
    the 'www' is also made for communicating, sharing, and exchanging, right?
    thank you for having the courtesy to respond to those who try to help you
    -1
    1. Besdu06
       
      no, it is not a protected site; it is a site where you can freely download documents, but you have to do it one by one, and there are more than 1500 of them. So the objective of httrack is to be able to retrieve these documents, which are either in .doc, .pdf, or HTML format.

      Thank you for your help^^
      0