How to Pull Historical Data from MT5 Using Python

The MetaTrader5 Python library will hand you a clean-looking dataframe without telling you it silently dropped the one column your cost model actually needed.


Why pull it yourself instead of exporting from the terminal

MT5’s terminal has a manual export option buried in the History Center, and for a quick look at a chart it’s fine. It’s a bad foundation for anything you’re going to backtest seriously, because it forces you into whatever timeframe and date range you clicked through the UI to select, it doesn’t fit into a repeatable pipeline, and it gives you no programmatic way to refresh the dataset as new bars print. Pulling data through the MetaTrader5 Python package instead means the exact same fetch can run every morning via a scheduled task, feed straight into whatever walk-forward or Monte Carlo step comes next, and produce an identical dataframe shape every time. The script below is a working version of that fetch — nothing exotic, which is exactly why it’s worth going through carefully. Most of what can go wrong here doesn’t throw an error. It just quietly gives you slightly wrong data that looks fine until a backtest built on it stops matching live behavior.

import MetaTrader5 as mt5
import pandas as pd
import logging
from datetime import datetime

logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(message)s')
logger = logging.getLogger(__name__)

MT5_SYMBOL = "USDJPY"
MT5_TIMEFRAME = mt5.TIMEFRAME_M15
BARS_TO_REQUEST = 35000
OUTPUT_CSV = f"{MT5_SYMBOL}_M15_data.csv"

def initialize_mt5():
    if not mt5.initialize():
        logger.error("MT5 Init Failed: %s", mt5.last_error())
        return False
    return True

def shutdown_mt5():
    mt5.shutdown()

def get_mt5_data(symbol, timeframe, n_bars):
    rates = mt5.copy_rates_from_pos(symbol, timeframe, 0, n_bars)
    if rates is None or len(rates) == 0:
        return None

    df = pd.DataFrame(rates)
    df['time'] = pd.to_datetime(df['time'], unit='s', utc=True)
    df.set_index('time', inplace=True)
    df = df[~df.index.duplicated(keep='first')]

    cols = ['open', 'high', 'low', 'close']
    if 'tick_volume' in df.columns:
        df.rename(columns={'tick_volume': 'volume'}, inplace=True)
    else:
        df['volume'] = 1.0

    return df[cols + ['volume']].copy()

Any symbol, any timeframe — three variables, not a rewrite

Nothing about this script is hardcoded to USDJPY specifically. MT5_SYMBOL, MT5_TIMEFRAME, and BARS_TO_REQUEST are the only three values that need to change to pull a different instrument, a different bar size, or a longer history — swap "USDJPY" for "EURUSD", "XAUUSD", or whatever else you want to look at, and the fetch, cleanup, and CSV export all run identically underneath. MT5_TIMEFRAME accepts any of the constants MetaTrader5 exposes — mt5.TIMEFRAME_M1, mt5.TIMEFRAME_H1, mt5.TIMEFRAME_D1, and so on — so the same function works whether you’re pulling M15 bars for a session-scoped pattern or H1 bars for something with a longer holding period.

The one thing worth checking before assuming the symbol swap “just works”: the string has to match exactly what your specific broker lists in Market Watch, which isn’t always the plain pair name. Some brokers append a suffix — EURUSD.raw, EURUSD.pro, occasionally something tied to account type — and copy_rates_from_pos will simply return None for a symbol string that doesn’t match, rather than suggesting the close alternative it probably meant. If a symbol swap comes back empty, checking the exact string in the terminal’s Market Watch panel is the first thing to rule out, before assuming the data doesn’t exist for that instrument at all.

initialize() is checking more than “is MT5 open”

mt5.initialize() returning False gets treated as a generic failure most of the time, but the causes split into a few distinct categories worth distinguishing. It fails if the terminal isn’t running, obviously, but also if the terminal is running under a different Windows user session than the Python process — a common VPS trap where the terminal was launched interactively but the script runs as a scheduled task under a service account. It also fails if path autodetection can’t find the installation, which happens more than expected when multiple broker-branded MT5 builds sit side by side on the same machine. mt5.last_error() differentiates these cases in its error code, so logging only the boolean return, rather than the full error tuple, throws away the one piece of information that tells you which of these you’re actually dealing with.

Bar count is not the same as calendar time

Requesting 35,000 M15 bars feels like it should map cleanly onto some number of months, but forex history has gaps built into it that make bar count and calendar time diverge in a way worth doing the arithmetic on before you assume you know your date range. A 15-minute timeframe produces 96 bars per trading day if the market ran continuously, but the weekly close from Friday evening to Sunday evening removes two days a week, and the actual number of bars per week ends up closer to 480 than the naive 672 a full seven-day week would suggest.

Assumption Bars
96 bars/day × 5 trading days 480/week
480 × ~52 weeks ~24,960/year
35,000 bars requested ~1.4 years of history

So a request for 35,000 bars is buying roughly a year and five months, not the “35,000 divided by bars-per-day” figure a quick mental calculation might produce if you forget to subtract weekends. This matters directly for anything downstream that cares about sample size — if your validation process needs a specific number of trades or a specific number of walk-forward windows, working backward from bar count to actual trading days, rather than calendar days, is the only way to know whether the history you just pulled is actually long enough for what you’re about to do with it.

The timezone label that isn’t telling the truth

This is the part of the script most likely to cause a problem nobody notices until much later. pd.to_datetime(df['time'], unit='s', utc=True) takes the Unix timestamp MT5 returns and tags the resulting datetime as UTC. But tagging a timestamp as UTC doesn’t verify that the underlying clock the timestamp was generated from was actually UTC — it just attaches a label. MT5’s bar timestamps are generated from the broker server’s clock, and broker server time is very often offset from true UTC by two, three, or more hours, with the exact offset frequently shifting on its own daylight saving schedule that doesn’t track any particular country’s DST rules consistently.

The practical effect: after this line runs, the dataframe’s index looks like proper UTC-aware timestamps, complete with the +00:00 offset markings that make it look verified. It isn’t. A bar timestamped 13:00 UTC in this dataframe might actually represent a candle that opened at 11:00 or 15:00 true UTC, depending on your specific broker’s server offset at that moment in the year. If you’re filtering this data by session boundaries — say, restricting analysis to the London/New York overlap, 13:00–16:00 UTC — using this dataframe’s index directly will silently pull the wrong bars, offset by however many hours your broker’s server clock differs from true UTC. The fix is to determine your specific broker’s UTC offset (visible in the terminal, or inferable by comparing a known news-release timestamp against the bar it lands in) and apply that correction explicitly before trusting the index for anything session-dependent. This is the exact same failure mode that causes an EA to appear to trade outside its intended hours — it just shows up here at the data-preparation stage instead of at execution time.

tick_volume is a proxy, not a measurement

The script renames tick_volume to volume, which is reasonable naming but worth not forgetting the substance of once the rename has happened. Forex is traded over the counter across many liquidity providers with no central exchange, so there’s no true consolidated volume figure the way there is for a listed equity or futures contract. tick_volume counts the number of price changes the broker’s feed recorded within the bar — a proxy for activity, not a count of contracts or lots actually traded. Two different brokers, feeding the same currency pair from different liquidity providers, will report different tick volume figures for the identical historical bar, because they’re counting quote changes from different sources, not measuring the same underlying thing.

This matters if your strategy or any regime-shift diagnostic leans on volume as a liquidity indicator, particularly across sessions. A quiet-looking tick volume reading during the Asian session (00:00–08:00 UTC) partly reflects genuinely thinner participation, but it also partly reflects that specific broker’s specific feed producing fewer quote updates during those hours — a mechanical property of that broker’s price stream, not a pure market-liquidity signal you could compare directly against another broker’s numbers or against actual traded volume from a different asset class entirely.

The column this script throws away

copy_rates_from_pos actually returns more fields than this script keeps. The underlying structured array includes a spread field alongside open, high, low, close, and tick_volume — a representative spread value for each bar, in points, at the broker’s own valuation. The script’s column selection, cols + ['volume'], drops it on the floor before it ever reaches the CSV.

This is worth fixing before this data goes anywhere near a backtest, because it’s precisely the information that determines whether the cost side of a backtest is honest or optimistic. A backtest that applies one flat, assumed spread across the whole dataset — the default behavior if this column isn’t captured — will systematically understate costs during the exact bars where spread was actually elevated: thin Asian-session stretches, the run-up to major news releases, and the illiquid pockets around weekly close and open. Adding 'spread' to the cols list costs nothing and turns this from a dataset that assumes a constant cost environment into one that at least has the broker’s own per-bar record of what costs actually looked like at the time.

Where the duplicate index rows actually come from

The deduplication line, df[~df.index.duplicated(keep='first')], is defensive code against a real, if infrequent, problem: a symbol’s history occasionally contains genuinely duplicated timestamps, most often around a daylight saving transition where the broker’s clock produces an ambiguous or repeated hour, or after the terminal’s local history cache has been rebuilt from more than one underlying data source and the stitch point overlaps by a bar or two. It’s cheap insurance to leave in, but it’s worth knowing it’s there for a reason rather than assuming it’s unnecessary boilerplate — a dataset that silently contained a duplicated bar would double-count that timestamp in any downstream aggregation without raising any error at all.

What this feeds into

Once this CSV exists, it’s the raw material for everything else in the validation pipeline — the train/test split that separates the history into a fitting window and a holdout window, the walk-forward chain that re-optimizes across sequential slices of it, and eventually the JSON strategy configs, with their start_hour, end_hour, sl, and tp fields, that get tested against it. Every one of those steps inherits whatever’s wrong with this file uncorrected. A timezone offset that’s off by three hours doesn’t announce itself later — it just quietly misaligns every session-boundary filter built on top of it, which is exactly why it’s worth getting this stage right before building anything on top of it, rather than treating data collection as the boring part that doesn’t need scrutiny.