How to download financial data with R

Hadrien Puche

Any financial analysis starts with data. Whether you want to analyze a stock, build a portfolio, measure risk, create a valuation model or develop trading strategies, the first step is always the same: obtaining financial data.

You could download data manually from websites such as Yahoo! Finance or Investing.com, but this quickly becomes tedious and time-consuming. It also limits the amount of data you can work with.

R allows us to automate this process and retrieve large amounts of financial information in just a few lines of code.

In this article, Hadrien PUCHE (ESSEC Business School, Grande École Program, Master in Management, 2023-2027) will help you to:

  • Download historical stock prices and market indices with R
  • Explore, clean, and visualize xts time-series data
  • Compute basic statistics and historical distributions
  • Compare multiple securities
  • Build the foundation needed for more advanced financial analysis

But first, what financial data can we actually download?

Financial professionals use many different categories of data across individual assets as well as portfolios and funds.

Market data (Easily downloadable for free via Yahoo! Finance)

  • Individual asset prices (e.g., individual stocks, corporate bonds)
  • Portfolios and funds (e.g., ETFs, mutual funds)
  • Currency exchange rates (e.g., EUR/USD)
  • Market indices (e.g., S&P 500)

Macroeconomic data (Available via the St. Louis Fed – FRED)

  • Inflation and Consumer Price Index (CPI)
  • Interest rates and bond yields
  • GDP growth and unemployment

Not all data sources are freely available. Many professional investors rely on paid platforms such as Bloomberg or FactSet to access fundamental accounting data (revenue, cash flows) and alternative data (satellite imagery, sentiment). Fortunately, market prices and macroeconomic indicators can easily be accessed for free using R for research and learning purposes.

While you can download macroeconomic data using specialized packages like fredr, we will keep things simple in this article and focus purely on extracting and modeling market prices using the open-source quantmod package.

A step-by-step guide

Follow the next steps to download your first financial dataset with R.

Step 1: Instal the required packages

If you have not yet installed R, refer to the setup guide published earlier in this series to configure your execution environment (RStudio).

Once your environment is ready, install the required packages by running this in your console (you only need to do this once):

install.packages(c("quantmod", "PerformanceAnalytics"))

Here is what these packages do:

  • quantmod: Short for Quantitative Financial Modelling Framework, it is a widely used R package for downloading and analyzing financial market data, including data from Yahoo! Finance.
  • xts: Short for eXtensible Time Series, this package is automatically installed with quantmod. It provides data structures specifically designed for time-indexed data.
  • PerformanceAnalytics: A package of econometric functions used to calculate returns and risk metrics.

Step 2: Import your R packages

Most financial analysis scripts begin by loading the packages required for the analysis using the library() function.

Create a new R script file. You can also save it wherever you want. Paste the following script and run it (as a reminder, you need to select the code you want with your mouse before running it):

library(quantmod)
library(PerformanceAnalytics)

Step 3: Download your first data

To download financial data, we use ticker symbols, which identify securities or market instruments within a given exchange or data provider. For example, AAPL represents Apple.

We will use the getSymbols() function. By passing arguments to the function, we can customize the output. Setting auto.assign = FALSE assigns the dataset directly to a variable that we can name aapl_data.

Let us download daily data for Apple’s stock price over the last five years (from 01/01/2019 to 31/12/2024 at that time)t.

# Download historical Apple stock data
aapl_data <- getSymbols("AAPL", src = "yahoo", from = "2019-01-01", to = "2024-01-01", auto.assign = FALSE)
 
# Display the first 5 rows
head(aapl_data, 5)

You will obtain this xts table with the following columns:

  • Open: opening price at the beginning of the trading day
  • High: highest price during the trading day
  • Low: lowest price during the trading day
  • Close: closing price at the end of the trading day
  • Volume: transaction volume during the trading day
  • Adjusted: closing price adjusted for stock splits and dividends

A screenshot from RStudio showing the output table of the getSymbols query

To keep things simple for this guide, we will focus strictly on the raw Close price. quantmod provides a convenient helper function called Cl() that instantly extracts just the closing price column from the dataset.

# Extract only the closing price
aapl_close <- Cl(aapl_data)
head(aapl_close, 3)

Add this code to your script, then highlight it with your mouse, and press run. Your RStudio should now display this:

A screenshot from RStudio showing the new output

Step 4: Use more precise queries for historical context

As financial analysts, we routinely extract specific timeframes to understand how assets behave under macroeconomic stress. Because our data is stored as an xts object, R makes it incredibly easy to slice time-series data using date ranges.

For example, examining the COVID-19 market shock in early 2020 provides a useful illustration of extreme market volatility. Let’s isolate Apple’s stock specifically during the COVID-19 market shock and initial recovery (January to June 2020):

# Isolate the COVID-19 crash using xts date subsetting (YYYY-MM-DD/YYYY-MM-DD)
covid_crash <- aapl_close["2020-01-01/2020-06-30"]
 
# Plot the isolated data
plot(covid_crash, main = "AAPL Stock Price - COVID-19 Crash & Recovery", col = "red", lwd = 2)

The output of the previous code cell showing the COVID crash

Step 5: Download the data for multiple stocks at the same time

Downloading multiple stocks is necessary for financial analysis that often requires comparing securities, analyzing sectors, or building a portfolio. Instead of issuing separate requests and risking misaligned dates, we can fetch all tickers at once.

Let’s download the data for six of the largest US banks: JPMorgan Chase, Bank of America, Wells Fargo, Citigroup, Goldman Sachs, and Morgan Stanley. Analyzing this sector is a classic way to measure the impact of interest rates on the broader economy.

# Define the major US bank tickers
bank_tickers <- c("JPM", "BAC", "WFC", "C", "GS", "MS")
 
# Download data into the global environment
getSymbols(bank_tickers, src = "yahoo", from = "2019-01-01", to = "2024-01-01")
 
# Extract only the closing prices and merge them into a single matrix
bank_prices <- merge(Cl(JPM), Cl(BAC), Cl(WFC), Cl(C), Cl(GS), Cl(MS))
 
head(bank_prices, 3)

The result is an xts object in which each column represents a stock and each row corresponds to a trading date.

Screenshot of the output of the previous cell showing the US Banks matrix

This table format is ideal for portfolio analysis and benchmarking. To save it for external use, you can export it as a CSV file:

# Save the data frame as a CSV file
write.csv(as.data.frame(bank_prices), file = "us_banks_data.csv")

Now that you have successfully downloaded your financial data, let’s see how you can clean it and then use it.

Inspecting and cleaning the dataset

Financial datasets may contain missing values (NA) for various reasons, including trading suspensions, differences in trading calendars, listing dates, or data-provider issues. Missing observations should be identified before computing returns or risk measures, as they may affect subsequent calculations. For this introductory example, we simply remove rows containing missing values using na.omit(). In applied financial analysis, however, the appropriate treatment depends on the source of the missing data and the objective of the analysis.

In R, we can easily remove any rows containing missing data using the na.omit() function.

# Check for missing values (returns the total count)
sum(is.na(aapl_close))

# Clean missing values by dropping rows with NAs
aapl_close <- na.omit(aapl_close)

# View the last few rows of the cleaned data
tail(aapl_close, 5)

A screenshot from RStudio showing the output of the tail function

Vizualizing your data

Let’s create our first chart to visualize the evolution of Apple’s stock price using the chartSeries() function, which is built specifically for financial time series.

chartSeries(aapl_close, 
            name = "Apple Stock Price", 
            theme = chartTheme("white"), 
            TA = NULL) # TA = NULL removes technical indicators for a clean chart

A screenshot of RStudio with the stock price visualization chart output

We have now:

  • Downloaded market data from Yahoo! Finance
  • Extracted the closing price and cleaned the data
  • Created a time-series plot to visualize stock prices

These core steps form the basis of empirical financial research and quantitative models.

Computing basic statistics and historical distributions

To evaluate stock performance and risk, we compute basic descriptive statistics. First, we calculate daily returns using the Return.calculate() function from the PerformanceAnalytics package.

# Calculate daily percentage returns (and remove the first NA row)
aapl_returns <- Return.calculate(aapl_close)
aapl_returns <- na.omit(aapl_returns)
 
# Compute summary statistics
mean_return <- mean(aapl_returns)
volatility <- sd(aapl_returns)
skew <- skewness(aapl_returns)
kurt <- kurtosis(aapl_returns)

print(paste("Mean Daily Return:", round(mean_return, 5)))
print(paste("Daily Volatility (Std Dev):", round(volatility, 4)))

Plotting historical distributions

Histograms display the frequency distribution of daily returns, helping us inspect distribution symmetry and tail risks.

# Return distribution histogram
hist(aapl_returns, breaks = 50, col = "salmon", main = "Historical Daily Return Distribution", xlab = "Daily Return")

The output of the previous cell – distribution histogram

Normalizing stock prices and computing returns

All stocks have different nominal prices. If Tesla trades at $350 and Nvidia at $220, it does not mean that Tesla performed better. To establish an accurate comparison, we execute two fundamental computations:

  1. Price normalization: We normalize all historical time series to a base index of 100, ensuring a standardized starting point.
  2. Return calculation: We compute periodic returns to measure performance independently of the nominal price level.
# Clean any missing data
bank_prices <- na.omit(bank_prices)

# Harmonize prices to Base 100 (Divide every row by the first row, multiply by 100)
normalized <- sweep(bank_prices, MARGIN = 2, STATS = as.numeric(bank_prices[1,]), FUN = "/") * 100
 
# Plot the performance comparison
plot(normalized, legend.loc = "topleft", main = "Performance Comparison (Base = 100)", ylab = "Growth of $100")

The output of the previous cell showing normalized prices

We can also compute daily returns across all stocks in one line:

bank_returns <- na.omit(Return.calculate(bank_prices))
head(bank_returns, 3)

This is a standard technique used by portfolio managers and equity analysts to compare growth trajectories.

Common pitfalls

When working with market data in R, beginners often run into the same issues:

  • Using the wrong ticker symbol
  • Comparing stocks without normalizing prices (Base 100)
  • Forgetting that markets are closed on weekends and holidays
  • Failing to clean and handle missing values (NA) using na.omit()
  • Using raw closing prices when adjusted prices are required: for long-term performance analysis, adjusted prices are generally preferable because they account for stock splits and dividends.

Overall, don’t forget to always inspect and clean your data before starting your analysis.

Exercises

Exercise 1: Basic data retrieval and price visualization (RACE)

Ferrari N.V. (RACE) presents an interesting case study in market dynamics: it is a car manufacturer that acts as a high-end luxury franchise. Its deliberate production scarcity, multi-year order backlogs, and immense pricing power decouple it from typical automotive boom-and-bust cycles.

Using the ticker symbol RACE, download the last five years of daily market data.

Your tasks:

  • Use the appropriate R functions to display the first 5 rows and the last 5 rows of the dataset to verify data integrity.
  • Extract the closing price and generate a line chart plotting the price over the entire 5-year period. Observe how its price trajectory reflects Ferrari’s distinctive positioning at the intersection of the automotive and luxury industries.

Exercise 2: time-series extraction and volume analysis on Tesla (TSLA)

Tesla is renowned for its high historical volatility and massive retail trading interest. The 2022-2024 window was particularly eventful for growth and electric vehicle stocks, marked by shifting supply chains and a rapid rise in interest rates.

Using the ticker symbol TSLA, extract the market data for the precise calendar period from January 1, 2022, to December 31, 2024 (using the from and to parameters).

Your tasks:

  • Identify the peak (highest closing price) and the trough (lowest closing price) over this period using the max() and min() functions.
  • Extract the Volume column (using Vo()) and calculate the average daily trading volume, a fundamental metric used by analysts to assess market liquidity.

Exercise 3: Comparative performance and risk profiling on Chinese tech companies

Chinese technology stocks often experience unique market cycles driven by distinct domestic regulatory environments and macroeconomic factors. Using the last five years of daily market data, compare the performance and risk characteristics of the following three US-listed ADRs:

  • Alibaba (BABA)
  • Baidu (BIDU)
  • PDD Holdings (PDD)

Questions:

  1. Which stock achieved the highest total cumulative return?
  2. Which stock was most volatile (highest standard deviation of daily returns)?
  3. Which one offered the best risk-adjusted profile over the period? Use the harpe ratio or any risk-adjusted measure

Download the solutions

To help you check your work and experiment further, you can download the complete R Script containing the full code, charts, and commentary for all exercises.

Download Solutions (.R Script)

What’s next?

Now that you know how to download financial data, compute returns, and analyze basic performance and risk measures in R, you are ready to delve into more advanced quantitative and corporate finance topics.

If you want to learn more about other programming languages, check out these two articles to learn how to install Python on your computer and use it to download financial data:

   ▶ Hadrien PUCHE How to Install and Run Python on Your Computer (A Step-by-Step Guide)

   ▶ Hadrien PUCHE How to download and model financial data with Python

About the Author

This article was written in September 2026 by Hadrien PUCHE (ESSEC Business School, Grande École Program, Master in Management, 2023-2027).

   ▶ Discover all posts by Hadrien PUCHE

How to install and run R on your computer (A step-by-step guide)

Hadrien Puche

Understanding and writing R code can be a valuable skill for your career. While general-purpose programming languages such as Python are more widely used, R is particularly well suited to statistical analysis, data visualization, and quantitative research.

In this article, Hadrien Puche (ESSEC Business School, Grande École Program, Master in Management, 2023-2027) will help you to:

  • Understand the core components of the R statistical ecosystem for finance
  • Compare development setups (RStudio Desktop vs. Visual Studio Code)
  • Install R alongside essential build tools (RTools / Xcode)
  • Set up RStudio Desktop as a purpose-built workspace
  • Install econometric packages via CRAN (like quantmod)
  • Run a test script

But first, what is R exactly?

Historically, R was created in 1993 as an open-source implementation of the S language, developed at Bell Labs for statistical computing.

Setting up your R workspace can be straightforward. We will rely on CRAN (the Comprehensive R Archive Network), R’s main public repository for packages, to install the packages required for our analysis and their dependencies. Let’s walk through deploying a professional quantitative workspace for R.

Quick vocabulary for beginners

Before we dive in, let’s define a few technical terms you will encounter frequently:

  • Package: A collection of reusable R functions, data, and documentation designed for a specific purpose. For example, quantmod provides tools for quantitative financial analysis and financial data retrieval.
  • Library: A directory on your computer, where your installed packages are stored. You will use the library() command in your code to load them. While developers often use the terms package and library interchangeably, technically you install a package into your library.
  • Dependency: A package that another package requires in order to work properly. R manages these dependencies automatically, so you do not have to take care of them, but do not be surprised if R installs many more packages than you initially requested.
  • Build Tools (Rtools / Xcode): Background software required by your computer to translate (or “compile”) raw source code into executable instructions. R frequently compiles financial packages directly on your machine, making these essential to prevent errors.

Choosing your development environment: RStudio or Visual Studio Code?

You generally have two choices when it comes to writing R code: RStudio Desktop and Visual Studio Code (VS Code).

  • RStudio Desktop: An Integrated Development Environment (IDE) built specifically for R. It features a 4-pane layout that lets you simultaneously view your scripts, console, environment variables (data frames loaded in memory), and charts.
  • Visual Studio Code: VS Code is a highly versatile code editor. You can run R in VS Code by installing the R extension and configuring the required R packages. This is a good choice if you plan to mix multiple programming languages in the same project, though configuring VS Code for R requires a bit more effort than RStudio.

RStudio Desktop 4-pane layout
RStudio Desktop Interface

Visual Studio Code running R
Visual Studio Code configured for R

For this guide, we will focus on setting up R and integrating it with RStudio, as it offers a purpose-built user experience for R.

Understanding the R architecture

The R architecture operates as follows:

  • Base R: The underlying computational engine that calculates the math and runs the logic.
  • Build tools (RTools / Xcode):oftware required to compile R packages from source when precompiled binary versions are not available. Most beginners will install packages from binaries, but having these tools available can prevent installation problems with packages that require compilation.
  • CRAN: The Comprehensive R Archive Network. This is the centralized, strictly regulated global repository for R packages.

Step-by-step installation guide

Step 1: Install R and build tools

First, we must install R. RStudio will not function without it.

  1. Go to the official CRAN Download Page.
  2. For Windows:
    • Click Download R for Windows > base > Download the latest R executable and install it using default settings.
    • Go back to the Windows page, click Rtools, and install the version matching your R installation. This is critical for compiling quantitative packages later.
  3. For macOS:
    • Click Download R for macOS and select the .pkg matching your chip (Apple Silicon or Intel).
    • To ensure packages compile correctly, open your Mac Terminal and run xcode-select --install to get the necessary developer tools.

Step 2: Install RStudio

Now, we install the integrated development environment (IDE) that we will use to write and execute R code.

  1. Head to the Posit RStudio Desktop website.
  2. Download the free version corresponding to your operating system (Windows or macOS).
  3. Run the installer. RStudio will normally detect the R installation completed in Step 1 automatically.

Step 3: Install packages from CRAN

Because R uses centralized package repositories such as CRAN, we can install the packages required for our financial analysis directly from the R console in RStudio.

  1. Launch RStudio.
  2. In the Console pane (bottom-left), type the following command and press Enter. This will reach out to CRAN and download the essential tools for market data and time-series analysis:

# Install quantmod for data retrieval, xts for time-series, and PerformanceAnalytics for risk metrics
install.packages(c("quantmod", "xts", "PerformanceAnalytics", "ggplot2"))

💡 Quick fix tip: R may occasionally ask whether you want to install a package from source when a binary version is also available. For beginners, the binary version is usually the simplest option. Installing from source may require Rtools on Windows or the Xcode Command Line Tools on macOS.

a screenshot of the output of the script when downloading the packages

Checking that everything is working as intended

Let’s verify your infrastructure by writing a short script that pulls actual market data.

  1. In RStudio, go to File > New File > R Script.
  2. Paste the following quantitative code into the top-left editor pane.
  3. Highlight all the text and press Ctrl+Enter (Windows) or Cmd+Enter (macOS) to run it.
# Load the quantitative financial modeling library
library(quantmod)
 
# Download historical financial data for Apple via Yahoo Finance API
getSymbols("AAPL", src = "yahoo", from = "2023-01-01", to = "2024-01-01", auto.assign = TRUE)
 
# Display the first 5 rows of the time-series array in the console
print(head(AAPL))
 
# Generate a financial chart with volume and Bollinger Bands for volatility analysis
chartSeries(AAPL, 
            name = "Apple Inc. (AAPL) Historical Prices", 
            theme = chartTheme("white"), 
            TA = c(addVo(), addBBands()))

A screenshot of RStudio after running the test script
After running the script, your RStudio should look like this

If the AAPL dataset appears in your top-right Environment pane, the data prints in your Console, and a professional candlestick chart renders in your bottom-right Plots pane, your R setup is fully operational.

Next steps & use cases

With R correctly configured, you are now equipped to tackle complex econometric and financial challenges. A good next step would be to learn how to download financial data with R.

You can learn how to do so in this article: How to download financial data with R

Useful resources

   ▶ CRAN (The Comprehensive R Archive Network): The main public repository for R packages, R distributions, and documentation. CRAN provides a centralized infrastructure for distributing and maintaining thousands of R packages used in statistical computing, econometrics, and quantitative research.

About the author

This article was written in September 2026 by Hadrien PUCHE (ESSEC Business School, Grande École Program, Master in Management, 2023-2027).

   ▶ Discover all posts by Hadrien PUCHE

Programming Languages for Quants

Jayati WALIA

In this article, Jayati WALIA (ESSEC Business School, Grande Ecole Program – Master in Management, 2019-2022) presents an overview of popular programming languages used in quantitative finance.

Introduction

Finance as an industry has always been very responsive to new technologies. The past decades have witnessed the inclusion of innovative technologies, platforms, mathematical models and sophisticated algorithms solve to finance problems. With tremendous data and money involved and low risk-tolerance, finance is becoming more and more technological and data science, blockchain and artificial intelligence are taking over major decision-making strategies by the power of high processing computer algorithms that enable us to analyze enormous data and run model simulations within nanoseconds with high precision.

This is exactly why programming is a skill which is increasingly in demand. Programming is needed to analyze financial data, compute financial prices (like options or structured products), estimate financial risk measures (like VaR) and test investment strategies, etc. Now we will see an overview of popular programming languages used in modelling and solving problems in the quantitative finance domain.

Python

Python is general purpose dynamic high level programming language (HLL). It’s effortless readability and straightforward syntax allows not just the concept to be expressed in relatively fewer lines of code but also makes it’s learning curve less steep.

Python possesses some excellent libraries for mathematical applications like statistics and quantitative functions such as numpy, scipy and scikit-learn along with the plethora of accessible open source libraries that add to its overall appeal. It supports multiple programming approaches such as object-oriented, functional, and procedural styles.

Python is most popular for data science, machine learning and AI applications. With data science becoming crucial in the financial services industry, it has consequently created an immense demand for Python, making it a programming language of top choice.

C++

The finance world has been dominated by C++ for valid reasons. C++ is one of the essential programming languages in the fintech industry owing to its execution speed. Developers can leverage C++ when they need to programme with advanced computations with low latency in order to process multiple functions fasters such as in High Frequency Trading (HFT) systems. This language offers code reusability (which is crucial in multiple complex quantitative finance projects) to programmers with a diverse library comprising of various tools to execute.

Java

Java is known for its reliability, security and logical architecture with its object-oriented programming to solve complicated problems in the finance domain. Java is heavily used in the sell-side operations of finance involving projects with complex infrastructures and exceptionally robust security demands to run on native as well as cross-platform tools. This language can help manage enormous sets of real-time data with the impeccable security in bookkeeping activity. Financial institutions, particularly investment banks, use Java and C# extensively for their entire trading architecture, including front-end trading interfaces, live data feeds and at times derivatives’ pricing.

R

R is an open source scripting language mostly used for statistical computing, data analytics and visualization along with scientific research and data science. R the most popular language among mathematical data miners, researchers, and statisticians. R runs and compiles on multiple platforms such as Unix, Windows and MacOS. However, it is not the easiest of languages to learn and uses command line scripting which may be complex to code for some.

Scala

Scala is a widely used programming language in banks with Morgan Stanley, Deutsche Bank, JP Morgan and HSBC are among many. Scala is particularly appropriate for banks’ front office engineering needs requiring functional programming (programs using only pure functions that are functions that always return an immutable result). Scala provides support for both object-oriented and functional programming. It is a powerful language with an elegant syntax.

Haskell and Julia

Haskell is a functional and general-purpose programming language with user-friendly syntax, and a wide collection of real-world libraries for user to develop the quant solving application using this language. The major advantage of Haskell is that it has high performance, is robust and is useful for modelling mathematical problems and programming language research.

Julia, on the other hand, is a dynamic language for technical computing. It is suitable for numerical computing, dynamic modelling, algorithmic trading, and risk analysis. It has a sophisticated compiler, numerical accuracy with precision along with a functional mathematical library. It also has a multiple dispatch functionality which can help define function behavior across various argument combinations. Julia communities also provide a powerful browser-based graphical notebook interface to code.

Related posts on the SimTrade blog

▶ Jayati WALIA Quantitative Finance

▶ Jayati WALIA Quantitative Risk Management

▶ Jayati WALIA Value at Risk

▶ Akshit GUPTA The Black-Scholes-Merton model

Useful Resources

Websites

QuantInsti Python for Trading

Bankers by Day Programming languages in FinTech

Julia Computing Julia for Finance

R Examples R Basics

About the author

The article was written in October 2021 by Jayati WALIA (ESSEC Business School, Grande Ecole Program – Master in Management, 2019-2022).