# Establish baseline performance

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/tune-network-workloads-on-bare-metal/)
- [Set up Tomcat](https://learn.arm.com/learning-paths/servers-and-cloud-computing/tune-network-workloads-on-bare-metal/1_setup/)
- [Establish baseline performance](https://learn.arm.com/learning-paths/servers-and-cloud-computing/tune-network-workloads-on-bare-metal/2_baseline/)
- [Tune performance with NIC queue counts](https://learn.arm.com/learning-paths/servers-and-cloud-computing/tune-network-workloads-on-bare-metal/3_nic-queue/)
- [NUMA-based tuning](https://learn.arm.com/learning-paths/servers-and-cloud-computing/tune-network-workloads-on-bare-metal/4_local-numa/)
- [IOMMU-based tuning](https://learn.arm.com/learning-paths/servers-and-cloud-computing/tune-network-workloads-on-bare-metal/5_iommu/)
- [Summary](https://learn.arm.com/learning-paths/servers-and-cloud-computing/tune-network-workloads-on-bare-metal/6_summary/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/tune-network-workloads-on-bare-metal/_next-steps/)

## Overview
In this section, you establish a baseline configuration before applying advanced techniques to tune the performance of Tomcat-based network workloads on an Arm Neoverse bare-metal instance.

> To avoid running out of file descriptors under load, raise the file‑descriptor limit on *both* the server and the client:
> 
> ```
> ulimit -n 65535
> ```

## Configure an optimal baseline before tuning
This baseline includes:
- Aligning IOMMU settings with Ubuntu defaults
- Setting a default CPU configuration
- Disabling access logging
- Setting optimal thread counts

## Align IOMMU settings with Ubuntu defaults
If you are using a cloud image (for example, AWS) with non-default kernel parameters, align IOMMU settings with the Ubuntu defaults: `iommu.strict=1` and `iommu.passthrough=0`.

Edit GRUB and add (or update) `GRUB_CMDLINE_LINUX`:
```
sudo vi /etc/default/grub
```
Add or update the line to include:
```
GRUB_CMDLINE_LINUX="iommu.strict=1 iommu.passthrough=0"
```
Update GRUB and reboot to apply the settings:
```
sudo update-grub && sudo reboot
```
Verify that the default settings have been successfully applied:
```
sudo dmesg | grep iommu
```
You should see that under the default configuration, `iommu.strict` is enabled, and `iommu.passthrough` is disabled:
```
__output__ [    0.877401] iommu: Default domain type: Translated (set via kernel command line)
__output__ [    0.877404] iommu: DMA domain TLB invalidation policy: strict mode (set via kernel command line)
__output__ ...
```

## Establish a baseline on Arm Neoverse bare-metal instances
To mirror a typical Tomcat deployment and simplify tuning, keep 8 CPU cores online and set the remaining cores offline. Adjust the CPU range to match your instance. The example below assumes 192 CPUs (as on AWS `c8g.metal-48xl`).

Set CPUs 8–191 offline:
```
for no in {8..191}; do sudo bash -c "echo 0 > /sys/devices/system/cpu/cpu${no}/online"; done
```
Confirm that CPUs `0–7` are online and the rest are offline:
```
lscpu
```
Example output:
```
__output__ Architecture:                aarch64
__output__   CPU op-mode(s):            64-bit
__output__   Byte Order:                Little Endian
__output__ CPU(s):                      192
__output__   On-line CPU(s) list:       0-7
__output__   Off-line CPU(s) list:      8-191
__output__ Vendor ID:                   ARM
__output__   Model name:                Neoverse-V2
__output__ ...
```
Restart Tomcat on the Arm instance:
```
~/apache-tomcat-11.0.10/bin/shutdown.sh 2>/dev/null
ulimit -n 65535 && ~/apache-tomcat-11.0.10/bin/startup.sh
```
From your `x86_64` benchmarking client, run `wrk2` (replace `<tomcat_ip>` with the server’s IP):
```
ulimit -n 65535 && wrk -c1280 -t128 -R500000 -d60 http://<tomcat_ip>:8080/examples/servlets/servlet/HelloWorldExample
```
Example result:
```
__output__ Thread Stats   Avg      Stdev     Max   +/- Stdev
__output__     Latency    16.76s     6.59s   27.56s    56.98%
__output__     Req/Sec     1.97k   165.05     2.33k    89.90%
__output__  14680146 requests in 1.00m, 7.62GB read
__output__  Socket errors: connect 1264, read 0, write 0, timeout 1748
__output__ Requests/sec: 244449.62
__output__ Transfer/sec:    129.90MB
```

## Disable access logging
Disabling access logs removes I/O overhead during benchmarking.

Edit `server.xml` and comment out (or remove) the **`org.apache.catalina.valves.AccessLogValve`** block:
```
vi ~/apache-tomcat-11.0.10/conf/server.xml
```
```
<!--
    <Valve className="org.apache.catalina.valves.AccessLogValve" directory="logs"
            prefix="localhost_access_log" suffix=".txt"
            pattern="%h %l %u %t &quot;%r&quot; %s %b" />
-->
```
Restart Tomcat:
```
~/apache-tomcat-11.0.10/bin/shutdown.sh 2>/dev/null
ulimit -n 65535 && ~/apache-tomcat-11.0.10/bin/startup.sh
```
Re-run `wrk2`:
```
ulimit -n 65535 && wrk -c1280 -t128 -R500000 -d60 http://<tomcat_ip>:8080/examples/servlets/servlet/HelloWorldExample
```
Example result:
```
__output__ Thread Stats   Avg      Stdev     Max   +/- Stdev
__output__     Latency    16.16s     6.45s   28.26s    57.85%
__output__     Req/Sec     2.16k     5.91     2.17k    77.50%
__output__  16291136 requests in 1.00m, 8.45GB read
__output__  Socket errors: connect 0, read 0, write 0, timeout 75
__output__ Requests/sec: 271675.12
__output__ Transfer/sec:    144.36MB
```

## Set optimal thread counts
To minimize contention and context switching, align Tomcat’s CPU‑intensive thread count with available CPU cores.

While `wrk2` is running, identify CPU‑intensive Tomcat threads:
```
top -H -p "$(pgrep -n java)"
```
Example output:
```
__output__ top - 08:57:29 up 20 min,  1 user,  load average: 4.17, 2.35, 1.22
__output__     Threads: 231 total,   8 running, 223 sleeping,   0 stopped,   0 zombie
__output__     %Cpu(s): 31.7 us, 20.2 sy,  0.0 ni, 31.0 id,  0.0 wa,  0.0 hi, 17.2 si,  0.0 st
__output__     MiB Mem : 386127.8 total, 380676.0 free,   4040.7 used,   2801.1 buff/cache
__output__     MiB Swap:      0.0 total,      0.0 free,      0.0 used. 382087.0 avail Mem
__output__     
__output__     PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
__output__    4677 ubuntu    20   0   36.0g   1.4g  24452 R  89.0   0.4   1:18.71 http-nio-8080-P
__output__    4685 ubuntu    20   0   36.0g   1.4g  24452 R   4.7   0.4   0:04.42 http-nio-8080-A
__output__    4893 ubuntu    20   0   36.0g   1.4g  24452 S   3.3   0.4   0:00.60 http-nio-8080-e
__output__    ...
```
You’ll typically see `http-nio-8080-e` and `http-nio-8080-P` threads as CPU-intensive. Because the `http-nio-8080-P` thread count is fixed at 1 (in current Tomcat releases), and you have 8 online CPU cores, set `http-nio-8080-e` to 7.

Edit `server.xml` and update the HTTP connector to set the worker thread counts and connection limits:
```
vi ~/apache-tomcat-11.0.10/conf/server.xml
```
Replace the existing connector:
```
<!-- Before -->
<Connector port="8080" protocol="HTTP/1.1"
           connectionTimeout="20000"
           redirectPort="8443" />
```
With the tuned settings:
```
<!-- After -->
<Connector port="8080" protocol="HTTP/1.1"
           connectionTimeout="20000"
           redirectPort="8443"
           minSpareThreads="7"
           maxThreads="7"
           maxKeepAliveRequests="500000"
           maxConnections="100000" />
```
Restart Tomcat and re-run `wrk2`:
```
~/apache-tomcat-11.0.10/bin/shutdown.sh 2>/dev/null
ulimit -n 65535 && ~/apache-tomcat-11.0.10/bin/startup.sh

ulimit -n 65535 && wrk -c1280 -t128 -R500000 -d60 http://<tomcat_ip>:8080/examples/servlets/servlet/HelloWorldExample
```
Example result:
```
__output__ Thread Stats   Avg      Stdev     Max   +/- Stdev
__output__     Latency    10.26s     4.55s   19.81s    62.51%
__output__     Req/Sec     2.86k    89.49     3.51k    77.06%
__output__  21458421 requests in 1.00m, 11.13GB read
__output__ Requests/sec: 357835.75
__output__ Transfer/sec:    190.08MB
```
With a solid baseline in place, you’re ready to proceed to NIC queue tuning, NUMA locality optimization, and IOMMU exploration in the next sections.
