# Install Apache Spark on Azure Cobalt 100 processors

## In this learning path

- [Introduction](https://learn.arm.com/learning-paths/servers-and-cloud-computing/spark-on-azure/)
- [Getting started with Microsoft Azure Cobalt 100, Azure Linux 3.0, and Apache Spark](https://learn.arm.com/learning-paths/servers-and-cloud-computing/spark-on-azure/background/)
- [Create an Azure Cobalt 100 Arm64 virtual machine](https://learn.arm.com/learning-paths/servers-and-cloud-computing/spark-on-azure/create-instance/)
- [Set up an Azure Linux 3.0 environment](https://learn.arm.com/learning-paths/servers-and-cloud-computing/spark-on-azure/container-setup/)
- [Install Apache Spark on Azure Cobalt 100 processors](https://learn.arm.com/learning-paths/servers-and-cloud-computing/spark-on-azure/deploy/)
- [Validate Apache Spark on Azure Cobalt 100 Arm64 VMs](https://learn.arm.com/learning-paths/servers-and-cloud-computing/spark-on-azure/baseline/)
- [Benchmark Apache Spark](https://learn.arm.com/learning-paths/servers-and-cloud-computing/spark-on-azure/benchmarking/)
- [Next Steps](https://learn.arm.com/learning-paths/servers-and-cloud-computing/spark-on-azure/_next-steps/)

Within your running docker container image or your custom Azure Linux VM, follow the instructions to install Spark.

Start by installing Java, Python, and other essential tools:

## Install Java, Python, and tools for Apache Spark

```bash
sudo tdnf update -y
sudo tdnf install -y java-17-openjdk java-17-openjdk-devel git maven wget nano curl unzip awk tar
sudo tdnf install -y python3 python3-pip
```

Verify Java installation:

```bash
java -version
```

The output will look like:

```
__output__
openjdk 17.0.16 2025-07-15 LTS
__output__
OpenJDK Runtime Environment Microsoft-11926147 (build 17.0.16+8-LTS)
__output__
OpenJDK 64-Bit Server VM Microsoft-11926147 (build 17.0.16+8-LTS, mixed mode, sharing)
```

Verify Python installation:

```bash
python3 --version
```

The output will look like:

```
__output__
Python 3.12.9
```

## Download and install Apache Spark on Azure Cobalt 100

You can now download and configure Apache Spark on your Arm-based machine:

```bash
wget https://downloads.apache.org/spark/spark-3.5.6/spark-3.5.6-bin-hadoop3.tgz
tar -xzf spark-3.5.6-bin-hadoop3.tgz
sudo mv spark-3.5.6-bin-hadoop3 /opt/spark
```

## Configure environment variables for Apache Spark

Add this line to `~/.bashrc` or `~/.zshrc` to make the change persistent across terminal sessions.

```bash
echo 'export SPARK_HOME=/opt/spark' >> ~/.bashrc
echo 'export PATH=$PATH:$SPARK_HOME/bin:$SPARK_HOME/sbin' >> ~/.bashrc
echo 'export JAVA_HOME=/usr/lib/jvm/msopenjdk-17/' >> ~/.bashrc
```

Apply changes immediately in your running shell:

```bash
source ~/.bashrc
```

## Verify Apache Spark installation on Azure Cobalt 100

```bash
spark-submit --version
```

You should see output like:

```
__output__
Welcome to
__output__
      ____
     / __/__  ___ _____/ /__
    _\ \/ _ \/ _ `/ __/  '_/
   /___/ .__/\_,_/_/ /_/\_\   version 3.5.6
      /_/
```

Spark installation is complete. You can now proceed with the baseline testing of Spark in the next section.
