Customize Your Crawl - Ryte Product Insights
How to Customize Your Crawl to Get the Best Results for Your Business
A comprehensive website analysis tailored to your needs is vital for sustainable website quality management. With Ryte, you can easily customize your analysis to get the best results for your business and derive appropriate optimization measures. This article describes the various configuration options and shows you how to adapt them to your needs.
Contents
- Introduction
- Use Cases
- Basic project settings
- Advanced project settings
- Previous analysis
- Test settings
- Conclusion
Introduction
Ryte uses its own crawler to analyze your website, helping you to identify issues that could harm your user experience. The crawling technology closely resembles Google’s crawler. It starts on the homepage of your project, and makes its way from page to page by following the internal link path. Just like the Google Crawler, the Ryte bot can be controlled. For example, you can instruct the crawler to exclude certain directories, pages or subdomains from the crawl, and therefore from your analysis. All Basic and Business Suite account owners benefit from the full functionality of the project settings.
For a quick tutorial about how to set up your project, check out this video.
The crawler settings can be found when clicking on your project settings on the top right hand corner throughout the Ryte Suite. The settings are divided into "Project setup" and "Advanced analysis". These can be modified individually for every project.
Figure 1: Access your project settings within the tool
Use cases
Different businesses have different requirements for a website analysis. The following use cases demonstrate how you can customize the crawler to suit your particular needs:
1: Analyze an entire website
You want to analyze the entire website, including all the sub-domains. This is recommended for smaller websites.
Crawl requirements: The entire website and all sub-domains should be included in the analysis.
2: Analyze a specific section of a website
As an SEO manager who works for a big company, you are responsible for the optimization of a specific directory and would like to check if all the pages in this directory are listed in the respective Sitemap.xml.
Crawl requirements: Only one directory should be crawled. Analyzing the entire website would be unnecessary and would not help identify the parts of your directory that need optimization.
3: Prepare for a website migration
You are planning to relaunch your website and would like to analyze the website’s performance and check for errors before its launch. For more advice about preparing for a relaunch, check out this article.
Crawl requirements: Website analysis and performance review despite htaccess password protection.
Basic project settings
You can make several basic adaptations to the analysis in the "project setup", or you can use the default settings. Let’s go through them step-by-step.
How many URLs should be analyzed?
Set the limit of the number of URLs that you want to crawl. Ideally, this should equal the number of URLs your website has. If you don’t know how many indexable URLs your website has, try the site query in Google: site:en.ryte.com. The number that appears on top will tell you how many of the domain’s pages are listed in the Google index – you can orient yourself towards it. The Ryte crawler is able to crawl from 100 to 21 million URLs.
Figure 2: How many URLs should be analyzed
How fast should the analysis be?
Here you can decide how fast your analysis should be. Parallel requests are used to determine the number of requests that should be sent by the crawler to the website. The more parallel requests you use, the faster your site will be analyzed. Large websites should set a high number of parallel requests to reduce the crawl duration. However, using more than 10 could cause your server to slow down, so check with your administrator or our support team for advice if you want to use more than 10 parallel requests.
Figure 3: Set the number of parallel requests
What should be analyzed?
We analyze certain aspects by default with our recommended crawl setting. However, you can untick or tick boxes as you require to ensure you analyze the data you need.
Figure 4: Recommended analysis settings
- Accept cookies: If your website uses cookies, here, you can instruct the crawler to accept cookies. This option is disabled by default, because this reveals issues for users (or crawlers) that do not accept cookies.
- Analyze images: The Ryte crawler regards pictures as autonomous resources and crawls them by default. If you only want to analyze your HTML content, you should untick it.
- Crawl subdomains: If your website has a lot of subdomains, you can crawl all subdomains by ticking this box. This option is activated by default.
- Obey robots.txt: You can instruct the Ryte crawler to regard or disregard the robots.txt.
- Analyse sitemaps: Would you like the crawler to download and analyze your sitemap.xml(s)?
- Schedule regular crawls: Click on "schedule analysis" to set a fixed interval for automatic crawls.
Figure 5: Schedule analysis
Advanced Analysis
How to analyze
In this section, you can specify more advanced settings. It is not necessary to change anything here if you want to use the default settings.
- Login data: Analyzing a website that is still in a testing environment is not a problem for our Ryte crawler. You can save login information, enabling you to crawl aspects of your website that are password protected.
Figure 6: Save login information for password-protected websites
- Robots.txt behavior: You can specify how the Ryte crawler should handle your website’s robots.txt file.
Figure 7: Robots.txt behavior in crawl settings
What to analyze
Homepage URL: The homepage URL determines the starting point of the Ryte crawler.
Analyze subfolder: This feature is useful if you want to optimize a specific directory on your website.
Analyze subdomains: For smaller websites, we would recommend crawling the entire website, including all subdomains.
Sitemap URLs: By default, the Ryte crawler analyzes the website’s sitemap.xml file that is saved under the default path www.domain.com/sitemap.xml.
Figure 8: Enter Sitemap.xml URLs
Ignore/include URLs
This function allows you to exclude URLs from the crawl. You can exclude entire directories, parameters, or individual pages from the crawl.
Previous analysis
All recently performed crawls within a project are listed in the "Previous analysis" tab. This table shows when the crawl started and finished, URLs found, analyzed URLs, and ignored URLs.
Figure 9: Previous analysis
Tip: If the number of found URLs is significantly higher than the number of crawled URLs, it would be a good idea to extend the Crawling Limit.
Test settings
Your crawler settings can be tested directly, live with the tab "Test Settings". We always recommend testing the crawl settings before starting a new crawl to ensure you get the data you require.
Check the following criteria:
- Status code of the website: Expected result: 200.
- Part of the project: Expected result: green check mark.
- Local links: A list with the website’s internal links should be displayed.
If all three criteria are met, you can rest assured that the crawl will be successful.
Figure 10: Test settings
Conclusion
Correct crawl settings help save time and effort. The various crawler settings make it easier for you to customize your analysis in the best way for your business. With customized crawl configurations, you can carry out your website analysis more efficiently and ensure that you are analyzing the data you need. You should test the settings before starting a crawl.