Skip to main content Explore View all products (200+) Microsoft Foundry Azure Copilot GitHub Copilot Azure Kubernetes Service (AKS) Azure Cosmos DB Azure Database for PostgreSQL Azure Arc Microsoft Fabric Linux virtual machines in Azure Foundry Models Foundry Agent Service Foundry IQ Foundry Tools Foundry Control Plane Observability in Foundry Control Plane Azure OpenAI in Foundry Models Azure Speech in Foundry Tools Azure Machine Learning View all databases Azure Cosmos DB Azure DocumentDB Azure SQL Azure Database for PostgreSQL Azure Managed Redis Microsoft Fabric Azure Databricks Linux virtual machines in Azure Windows Server on Azure Azure Functions Azure Virtual Machine Scale Sets Azure API Management Azure Container Apps Azure Kubernetes Service (AKS) Azure Kubernetes Fleet Manager Azure Container Registry Azure Red Hat OpenShift Azure Container Instances Azure Container Storage Azure Arc Azure Local Microsoft Defender for Cloud Azure Monitor Microsoft Sentinel Azure Migrate View all solutions (40+) Cloud solutions for small and medium businesses Cloud migration and modernization center Data analytics for AI Azure Databases AI apps and agents Microsoft Marketplace Microsoft Sovereign Cloud AI apps and agents Responsible AI with Azure AI Infrastructure Data analytics for AI Machine learning operations (MLOps) Low-code application development on Azure Integration Services Serverless computing DevOps Migration and modernization center .NET apps migration Databases on Azure Linux on Azure Oracle on Azure SAP on the Microsoft Cloud Adaptive cloud High-performance computing (HPC) Infrastructure as a service (IaaS) Resiliency Azure Essentials Frontier Accelerate for Azure FinOps on Azure Microsoft Marketplace Azure pricing overview Create an Azure account Free Azure services Flexible purchase options Pricing calculator FinOps on Azure Maximize ROI from AI Azure savings plans Azure reservations Azure Hybrid Benefit Virtual Machines Azure SQL Microsoft Foundry Microsoft Fabric Azure Kubernetes Service (AKS) Microsoft Defender for Cloud View more Software Development Companies Microsoft Marketplace Find a partner Resources for Azure partners Get started with Azure Customer stories Analyst reports, white papers, and e-books Videos Learn more about cloud computing Documentation Explore Azure portal Developer resources Quickstart templates Resources for startups Developer community Students Azure for partners Blog Events and Webinars Learn Support Contact Sales Get started with Azure Sign in

Today, we are pleased to announce that Apache Spark v1.6.1 for Azure HDInsight is generally available. Since we announced the public preview, Spark for HDInsight has gained rapid adoption and is now 50% of all new HDInsight clusters deployed. With GA, we are revealing improvements we’ve made to the service to make Spark hardened for the enterprise and easy for your users. This includes improvements to the availability, scalability, and productivity of our managed Spark service.

What is Apache Spark?

Apache Spark is an open source processing framework that runs large-scale data analytics applications in-memory. This allows Spark to deliver queries up to 100 times faster than traditional big data solutions, along with a common execution model for various tasks like extract-transform-load (ETL) processes, batch queries, interactive queries, real-time streaming, machine learning, and graph processing on data stored.

What is Microsoft’s Apache Spark for Azure HDInsight?

Microsoft has been on a journey to make big data easy and more approachable. This is encompassed in Cortana Intelligence, our big data and analytics suite. As part of this solution, we offer Azure HDInsight, Microsoft’s managed Hadoop and Spark cloud service that runs the Hortonworks Data Platform. Spark for Azure HDInsight offers customers an enterprise-ready Spark solution that’s fully managed, secured, and highly available and made simpler for users with compelling and interactive experiences.

clip_image001
  • Enterprise-ready Spark implementation: We’ve drawn on years of experience working with enterprise customers, running some of the largest data projects in the world. For Spark to run at this scale, Microsoft had to ensure that it is highly available, scalable, and secure.
    • For high availability, Microsoft worked with Hortonworks to add capabilities to the YARN resource manager and co-led “Project Livy” with Cloudera and other organizations to create an open source Apache licensed REST web service for managing long running Spark contexts and submitting Spark jobs. This new capability was designed to make Spark a more robust back-end for running interactive notebooks and allow other applications to leverage Spark for their interactive workloads. By ensuring high availability with Spark, we now offer the highest guarantee for Spark in the market with a 99.9% service level agreement.
    • To ensure that Spark will run at scale, we are announcing integration between Spark and Azure Data Lake Store. This will allow Spark to store and process data of any size built on a repository designed for the cloud to capture data of any size, type and speed without forcing changes to your application as data scales.
    • For securing Spark, we are enabling role-based data access at the storage level through the integration of Spark and Data Lake Store.
  • Spark made simpler: Our goal with big data is to make it accessible for everybody. With Spark for HDInsight, we have designed new productivity experiences for the different audiences that use Spark including the data engineer working on ETL jobs, the data scientists who are performing experimentation and the business analysts who are creating dashboards.
    • For the data engineer and developers, we introduced deep integration with the IntelliJ IDE. This allows developers to code with native authoring support for Scala and Java, local testing, remote debugging, and the ability to submit Spark applications to the Azure cloud.
    • For data scientists, we introduced out-of-the-box integration with Jupyter (iPython) notebooks allowing you to create narratives that combine code, statistical equations, and visualizations that tell a story about the data. This environment is ideal for extracting data from any source and iteratively building ML models while writing exploratory queries to visualize and understand properties of the data. We made this possible by working with the Jupyter OSS community to enhance the kernel to allow Spark execution through a REST endpoint. As a result, Jupyter notebooks are now accessible within HDInsight out-of-the-box.
    • For the business analysts, we offer integration with Power BI alongside other BI tools like Tableau, SAP Lumira, and QlikView. This lets you build interactive visualizations over data of any size. In addition to the traditional dashboards, Power BI offers a streaming connector that has integration with Spark allowing you to publish real-time events from Spark Streaming directly to Power BI.

How do I get started?

To get started, customers will need to have an Azure subscription or a free trial to Azure. With this in hand, you should be able to get a Spark cluster up and running in minutes by going through this getting started guide.

Also, head over to watch this Channel 9 video below on Azure Fridays:

clip_image002

Overview

Documentation and How-To’s:

More Resources

WE ARE MICROSOFT

Explore Microsoft Foundry

The future of AI starts here. Envision your next great AI app with the latest technologies. Get started with Azure.