Skipping links[edit source]

Why's it skipping some links while archiving other ones?. For examples, 39-40 or 44-48 in Sanskrit cinema or many such links in Dhee (Sanskrit Film). Dokania😎(📩) 09:26, 9 October 2026 (IST)Reply

Those links have no snapshots on archive.org. The bot adds only existing archives. Unnesssary regional news sites are often not crawled on archive.org. 🍁Bharatwiki Socrates🍁(📩) 09:51, 9 October 2026 (IST)Reply
I think it should archive the page if there's not one already. Regional news sites are most prone to link rot so they must be archived. Dokania😎(📩) 09:53, 9 October 2026 (IST)Reply
As I mentioned, the bot only adds archive links for URLs that already have snapshots on archive.org. Yes, archive.org does capture many sites, but not all. If a site isn't in their archive, no tool can create a link for it, this limitation exists on every encyclopedia. Also, please avoid requesting redirect pages. They have no citation URLs to archive, so the bot skips them automatically. 🍁Bharatwiki Socrates🍁(📩) 10:01, 9 October 2026 (IST)Reply
I know tools which archive new links as well, It's not very hard to integrate. Have you checked [1] Dokania😎(📩) 10:06, 9 October 2026 (IST)Reply
I think you misunderstood my comment. Yes, new websites can also be archived by this tool. However, if a website has been created recently, there is a possibility that Archive.org has not yet captured its pages, so they cannot be archived. That is a technical limitation.
Archive.org captures millions of websites daily. I don't know what you understand by archiving url, but it is not the tool itself that archives a website. Archive.org actually captures and stores the website's pages, while the tool simply retrieves the archived links. 🍁Bharatwiki Socrates🍁(📩) 10:12, 9 October 2026 (IST)Reply
@Dokania I think you have a high misunderstanding of how the Wayback Machine works. This tool can archive URLs regardless of whether the website is new or old if it was captured by wayback machine. However, if Archive.org has never captured the website's pages, the tool cannot retrieve an archived version of them. This is a technical limitation from archive.org, not a limitation of the tool itself. 🍁Bharatwiki Socrates🍁(📩) 10:18, 9 October 2026 (IST)Reply
Yes, I understand perfectly how it works. I am just saying that the bot should send a capture request to archive.org if it's unable to find any archived version of it. Dokania😎(📩) 10:22, 9 October 2026 (IST)Reply
Yes it does, here's how the bot retrieves archived URLs:
1. Parses the article's wikitext and extracts URLs from citation templates that lack an "archive-url" parameter.
2. Queries Archive.org's CDX API ("/cdx/search/cdx?url=<URL>&output=json"), which returns a timestamp and the original URL if a snapshot exists.
3. Constructs the archive link using "/web/<timestamp>/<url>".
4. Adds the "archive-url", "archive-date", and "url-status=live" parameters to the citation template.
5. Saves the changes via the MediaWiki API.
  • Which type of sites get skipped:
1. If the site is dead or deleted.
2. If it was never captured by wayback machine of archive.org.
🍁Bharatwiki Socrates🍁(📩) 10:36, 9 October 2026 (IST)Reply