Customize Your Crawl - Ryte Product Insights

How to Customize Your Crawl to Get the Best Results for Your Business

A comprehensive website analysis tailored to your needs is vital for sustainable website quality management. With Ryte, you can easily customize your analysis to get the best results for your business and derive appropriate optimization measures. This article describes the various configuration options and shows you how to adapt them to your needs.

Contents

  1. Introduction
  2. Use Cases
  3. Basic project settings
  4. Advanced project settings
  5. Previous analysis
  6. Test settings
  7. Conclusion

Introduction

Ryte uses its own crawler to analyze your website, helping you to identify issues that could harm your user experience. The crawling technology closely resembles Google’s crawler. It starts on the homepage of your project, and makes its way from page to page by following the internal link path. Just like the Google Crawler, the Ryte bot can be controlled. For example, you can instruct the crawler to exclude certain directories, pages or subdomains from the crawl, and therefore from your analysis. All Basic and Business Suite account owners benefit from the full functionality of the project settings.

For a quick tutorial about how to set up your project, check out this video.

The crawler settings can be found when clicking on your project settings on the top right hand corner throughout the Ryte Suite. The settings are divided into "Project setup" and "Advanced analysis". These can be modified individually for every project.

Figure 1: Access your project settings within the tool

Use cases

Different businesses have different requirements for a website analysis. The following use cases demonstrate how you can customize the crawler to suit your particular needs:

1: Analyze an entire website

You want to analyze the entire website, including all the sub-domains. This is recommended for smaller websites.

Crawl requirements: The entire website and all sub-domains should be included in the analysis.

2: Analyze a specific section of a website

As an SEO manager who works for a big company, you are responsible for the optimization of a specific directory and would like to check if all the pages in this directory are listed in the respective Sitemap.xml.

Crawl requirements: Only one directory should be crawled. Analyzing the entire website would be unnecessary and would not help identify the parts of your directory that need optimization.

3: Prepare for a website migration

You are planning to relaunch your website and would like to analyze the website’s performance and check for errors before its launch. For more advice about preparing for a relaunch, check out this article.

Crawl requirements: Website analysis and performance review despite htaccess password protection.

Basic project settings

You can make several basic adaptations to the analysis in the "project setup", or you can use the default settings. Let’s go through them step-by-step.

How many URLs should be analyzed?

Set the limit of the number of URLs that you want to crawl. Ideally, this should equal the number of URLs your website has. If you don’t know how many indexable URLs your website has, try the site query in Google: site:en.ryte.com. The number that appears on top will tell you how many of the domain’s pages are listed in the Google index – you can orient yourself towards it. The Ryte crawler is able to crawl from 100 to 21 million URLs.

Figure 2: How many URLs should be analyzed

How fast should the analysis be?

Here you can decide how fast your analysis should be. Parallel requests are used to determine the number of requests that should be sent by the crawler to the website. The more parallel requests you use, the faster your site will be analyzed. Large websites should set a high number of parallel requests to reduce the crawl duration. However, using more than 10 could cause your server to slow down, so check with your administrator or our support team for advice if you want to use more than 10 parallel requests.

Figure 3: Set the number of parallel requests

What should be analyzed?

We analyze certain aspects by default with our recommended crawl setting. However, you can untick or tick boxes as you require to ensure you analyze the data you need.

Figure 4: Recommended analysis settings

Figure 5: Schedule analysis

Advanced Analysis

How to analyze

In this section, you can specify more advanced settings. It is not necessary to change anything here if you want to use the default settings.

Figure 6: Save login information for password-protected websites

Figure 7: Robots.txt behavior in crawl settings

What to analyze

Homepage URL: The homepage URL determines the starting point of the Ryte crawler.

Figure 8: Enter Sitemap.xml URLs

Ignore/include URLs

This function allows you to exclude URLs from the crawl. You can exclude entire directories, parameters, or individual pages from the crawl.

Previous analysis

All recently performed crawls within a project are listed in the "Previous analysis" tab. This table shows when the crawl started and finished, URLs found, analyzed URLs, and ignored URLs.

Figure 9: Previous analysis

Tip: If the number of found URLs is significantly higher than the number of crawled URLs, it would be a good idea to extend the Crawling Limit.

Test settings

Your crawler settings can be tested directly, live with the tab "Test Settings". We always recommend testing the crawl settings before starting a new crawl to ensure you get the data you require.

Check the following criteria:

If all three criteria are met, you can rest assured that the crawl will be successful.

Figure 10: Test settings

Conclusion

Correct crawl settings help save time and effort. The various crawler settings make it easier for you to customize your analysis in the best way for your business. With customized crawl configurations, you can carry out your website analysis more efficiently and ensure that you are analyzing the data you need. You should test the settings before starting a crawl.