What downloading a website means and when you'd do it
Downloading a website means saving the HTML files, images, and other content from a web page to your computer so you can view it without an internet connection. This is different from saving a single image or PDF — you're capturing the whole page structure.
People download websites for several reasons: keeping a copy of information that might disappear, reading articles offline during travel, archiving research for a project, or preserving a page before it changes. The method you use depends on what you're trying to save and how much of the site you need.
Key Takeaways
- Your browser's built-in save function works for single pages but may not capture all images and linked content correctly.
- Dedicated download tools like HTTrack or Wget can mirror entire websites, including multiple pages and all assets, but require more setup.
- Saving as PDF preserves the page's appearance but loses interactivity and doesn't work well for multi-page sites.
- Some websites block automated downloading in their terms of service, so check before using tools that crawl multiple pages.
Using your browser to save a single page
The simplest method is built into every browser. In Chrome, Firefox, Safari, or Edge, press Ctrl+S (Windows) or Command+S (Mac), or go to File menu and select Save Page As. A dialog box will appear asking where to save the file and what format you want.
Choose "Webpage, Complete" or "HTML Only" depending on your browser. "Complete" saves the page as an HTML file plus a folder containing images and stylesheets, so the page looks right when you open it later. "HTML Only" saves just the text and structure, which takes less space but may look plain without the styling.
This method works well for single articles, reference pages, or documentation. It does not always capture content loaded by JavaScript (interactive elements that appear after the page loads), and it may miss some linked resources if the site uses unusual hosting.
Saving as PDF for offline reading
If you only need to read the page and don't care about interactivity, saving as PDF is faster and more reliable. Press Ctrl+P (Windows) or Command+P (Mac) to open the print dialog, then select "Save as PDF" instead of printing to paper.
PDF preserves the page's layout and fonts exactly as they appear on screen, and the file works on any device. The downside is that PDFs don't include clickable links the way the original page does, and they work poorly for very long pages or sites with multiple sections you'd need to navigate between.
Some websites format poorly when converted to PDF — text might overlap, images could be cut off, or the page might span dozens of PDF pages when printed. Test it on a small section first if you're unsure.
Downloading multiple pages with HTTrack or Wget
If you need to save an entire website or many linked pages, a dedicated download tool is necessary. HTTrack (Windows, Mac, Linux) and Wget (command line, all platforms) are the most common free options. Both crawl through a website following links and save everything they find to a folder on your computer.
HTTrack has a graphical interface: download it from httrack.com, launch it, enter the website URL, set how many levels deep you want to crawl (level 1 is just the homepage, level 3 includes pages linked from those pages), and click Start. It will create a folder with the entire site structure inside, which you can then open in your browser offline.
Wget is a command-line tool that does the same thing but requires typing commands in a terminal. It's more powerful for advanced users but steeper to learn. On Windows, you can download Wget from gnuwin32.sourceforge.net; on Mac and Linux, it's usually pre-installed or available through package managers.
Understanding website terms of service and robots.txt
Before downloading an entire website, check whether the site's terms of service allow it. Many sites prohibit automated downloading or crawling, especially news sites and social media platforms. Downloading against the terms of service could expose you to legal issues.
Most websites have a file called robots.txt in their root directory that tells crawlers which parts they can and cannot access. You can view it by typing the website URL followed by /robots.txt in your browser (for example, example.com/robots.txt). If it says "Disallow: /", the site does not want to be crawled.
Personal blogs, documentation sites, and educational resources are usually fine to download for personal use. Commercial sites, paywalled content, and copyrighted material require more caution. When in doubt, contact the site owner or assume downloading is not permitted.
Fixing broken links and missing images after download
After downloading a website, you may find that some images don't display or links don't work. This happens because the downloaded files reference the original website's server, and those references break when you view the files offline.
HTTrack usually handles this automatically by rewriting links to point to the local files instead. If you used your browser's save function and images are missing, the images may not have been saved to the folder. Open the folder where you saved the page, look for a subfolder named something like "example.com_files" or similar, and check whether image files are inside.
If images are truly missing, you can try downloading the page again using HTTrack, which is more thorough about capturing all assets. For a few missing images, you can also download them individually by right-clicking each broken image in your browser and selecting "Save image as."
Frequently Asked Questions
Can I download a website that requires a login?
Browser save functions will only capture what you see after logging in, so that works for single pages. HTTrack and Wget cannot log in automatically without additional configuration, so they won't reach password-protected content. For sites behind a login, your best option is saving individual pages through your browser as you navigate.
What's the difference between saving as HTML and saving as MHTML?
HTML saves the page as a text file plus a separate folder of images and styles. MHTML (also called MHT) bundles everything into a single file, which is easier to move around but larger and less compatible with older browsers. Most modern browsers support both, so choose based on whether you prefer one file or a folder structure.
Will downloading a website use a lot of storage space?
A single article typically takes 1 to 5 megabytes. An entire small website might be 50 to 500 megabytes. Large sites with many pages and high-resolution images can reach several gigabytes. Before downloading, set HTTrack to limit file size or depth so you don't accidentally download more than you need.
Can I re-upload a downloaded website to the internet?
No. Downloading a website for personal offline reading is generally permitted, but republishing it online violates copyright law. The original creator owns the content, and hosting it yourself without permission is infringement. The only exception is if the site is explicitly marked as free to reuse under a Creative Commons or similar license.
Why do some websites look different after I download them?
Websites often load content dynamically using JavaScript, which runs in your browser but doesn't always save to files. Downloaded pages may be missing animations, interactive features, or content that appears only after scrolling. This is a limitation of how downloads work, not a problem with your method.