# What is Recommendation Wizard for Microsoft Azure and how does it work?

**Cirrus Data Cloud (CDC)** is integrated with various storage providers to automatically allocate destination storage volumes that matches the migration source volumes. Unlike integrations with other storage providers, where users must specify additional volume properties during auto allocation process such as the performance tiers, throughput and IOPS, etc., Cirrus Data Cloud’s Microsoft Azure integration has been expanded to automatically recommend a volume type, size and performance parameters based on the real performance of the source volume.


Performance statistics and IO access information of the existing source volumes is crucial to accurately recommend destination volumes’ size and performance. When a user installs the **Cirrus Migrate Cloud (CMC)** on a Windows or Linux host, it immediately begins sampling IO statistics every second on all the attached volumes. CMC collects data about the IOPS, throughput, latency, types of IO (read/write/UNMAP), IO size, randomness, etc. These samples are then consolidated and downsampled by the minute. By default, after 30 days these IO statistics will be rotated to minimize the storage usage on the volume that CMC is installed on. All I/O statistics and access information are stored locally inside user’s hosts.


When a user starts the **Recommendation Wizard for Microsoft Azure**, they can specify various tunable parameters such as the amount of time CDC should consider when aggregating statistics from the source host. They can also specify whether they expect their performance and capacity requirements to be changed post-migration. Finally, they can select a recommendation profile among “Low Cost”, “Balanced” and “High Performance”. Once these parameters are customized (if needed) and the recommend button is clicked, CDC begins performing a series of tasks to recommend an accurate volume size and performance to the user.


Cirrus Data Cloud will first fetch an appropriate amount of data points for the source volume and aggregates them into min, max, average and standard deviation for each data type (IOPs, throughput, etc.). Then, the destination host which is running in Microsoft Azure will be queried for information that will be used to determine what disk types are eligible to be recommended. CMC uses the local metadata service along with Azure’s public api to query information about destination host. Information that may affect recommendation includes location, availability zones, machine type, ultra disk compatibility, shared disk compatibility, etc.. All of this information is gathered and then delivered to Cirrus Data Cloud’s recommendation service that is responsible for finding an appropriate recommendation and estimating the cost.


In Microsoft Azure, there are a number of disk types with different characteristics, ranging from Standard HDD to Ultra Disk. For purposes of calculation the disk types are split into two categories: *non-scaling* and *scaling* disk types.

* *Non-scaling* disk types such as Standard HDD and Premium SSD have performance tiers that are based on the capacity of the created disk. When creating these types of disks, performance characteristics are not customizable. Instead, users must create the disks with enough capacity to meet the minimum capacity of various performance tiers.
* *Scaling* disk types such as Premium SSD v2 and Ultra Disk are also allocated and performance-bounded based on capacity of the volumes. However, for each GiB allocated, users are allowed to provision a certain amount of maximum IOPS and for each IOPS provisioned users can allocate up to a certain amount of maximum throughput.

Cirrus Data Cloud’s Recommendation Wizard has full awareness of different disk types and disk tiers in Azure including minimum capacity, how fast IOPs/throughput scales and the performance characteristics of each disk tier.

While there are limited selections from the user-facing user interface, each of the aforementioned disk tiers and types is a unique SKU and each of them has a number of price meters associated with it. For example, a Premium SSD v2 disk has an independent meter for capacity, throughput and IOPS. The prices of these meters also varies across regions. When providing recommendation, Cirrus Data Cloud utilizes **Microsoft Azure Retail Rates Prices API** to gather the current up-to-date location-specific prices of all the price meters involved and take into account different pricing factors including the “free” IOPS/Throughput that comes with Premium SSD v2 and Ultra Disk.


With the gathered source volume statistics, destination host metadata and user configured parameters, Cirrus Data Cloud’s recommendation service will then generate a set of target values for each of the recommendation profiles that serve as a minimum when used in matching. The target values are then checked against every disk type and tier that are available to the selected destination host. If the disk type and tier can accommodate the target values, the recommendation service will then calculate the real performance values and the corresponding price.


For example, (only using Premium SSDs in this example), the target values determined were 80 GiB, 400 IOPs, 120 MB/sec. P6 Tier would not qualify as its maximum IOPS and throughput are too low. Instead, P10 and above would be considered and we would determine that P10 would be the cheapest option. Therefore, our finalized values used for recommendation would be the minimum disk capacity for P10 and the IOPs/throughput of P10. Even though the original source volume was only 80 GiB the recommendation service would recommend a 128 GiB capacity in order to achieve the needed performance.


The complexity involved in recommending the correct non-scaling disks is rather little: Cirrus Data Cloud’s recommendation service will check the performance bounds of every disk tier and find the lowest cost option.


However, scaling disks are much more sophisticated because capacity, IOPS and throughput can all be tweaked individually and each has a cost associated to it. In order to calculate the optimal parameters and consequently the price of the recommendation, the recommendation service works backwards. CDC understands that users get up to X IOPS per GiB allocated and up to Y throughput per an IOPS. Starting with the target value for throughput, CDC determines if the target IOPs is enough to satisfy the desired throughput. If not, it will adjust the final recommended IOPS and capacity and re-calculate. Once the recommendation is determined, it will use the pricing information and calculate the price of the real recommendation including free throughput/IOPS.


Every time the service performs a recommendation, it always returns the lowest price configuration that satisfies the recommendation target values. However, because different target values are generated based on the recommendation profile selected by the user, it may result in recommendations that are more/ less performant than needed or the costs are too high to particular users. If the result of the recommendation service is not desirable, users can tweak the settings such as expected growth and recommendation profile in order to get a more appropriate recommendation. Users may also override the recommendation and manually select the disk parameters. In those cases, Cirrus Data Cloud’s recommendation wizard will still perform cost estimation based on the manually entered parameters.


Once the recommendation is finalized and accepted, Cirrus Data Cloud will automatically allocate the volumes, map them to the destination volumes, and create migration sessions that have minimum impact to production workload.