Get in touch

9io.ai / Blog

Rate limits, holidays and atomic writes in our market data layer

How we fetch daily prices for US and Indian stocks without tripping Yahoo’s limits, how we tell an NSE holiday from a missing file, and what we haven’t fixed.

Key takeaways

  • Yahoo throttling often arrives as an empty result, so treat empty and all-NaN frames as rate limits.
  • Back off for minutes, carry the wait across chunks, and add jitter when clients share an address.
  • Cache an NSE holiday only after both file formats return 404 for a date at least two days old.
  • Don’t hand-type holiday calendars. When we checked ours against NSE’s 2026 circular, 7 of 15 dates were wrong.
  • Bhavcopy prices are unadjusted, so a split or bonus issue looks like a crash until you adjust it.

9io Alpha, our trading research platform, needs a year of daily bars for about 1,500 US stocks and for every equity on India’s National Stock Exchange (NSE). The two sources fail in different ways. Yahoo, which we reach through the open-source yfinance library, serves one stock per request and throttles without raising an error. NSE publishes one file per trading day for the whole market, and the hard part is telling a holiday from a file that isn’t there yet.

On the US side we run one sweep per process, in chunks of 100 symbols with 7-second pauses, and back off from 45 seconds to a 300-second cap when a chunk comes back empty. On the NSE side, a day counts as a holiday only when both file formats return 404 and the date is at least two days old. Every cache file is written under a temporary name and renamed into place.

This post covers both pipelines, the calendar mistakes we found while writing it, how the caches are written, and the split adjustment we still don’t make. Our post on look-ahead bias in our backtester covers what happens to these prices next. It describes engineering and isn’t investment advice.

Two sources that fail in different ways

US India
Source Yahoo, through yfinance NSE’s daily bhavcopy files
One download gets One stock’s history One day for the whole exchange
Universe S&P Composite 1500, about 1,500 stocks Every EQ-series stock in the file
First build of a year About 1,500 requests About 250 requests
Refresh Only stale or missing stocks About one request per new trading day
Stored as One file per stock One file per day
Main failure Throttling that returns empty data Holidays and late files both return 404
Liquidity floor before scanning Price of $3, $10M average daily traded value Price of ₹20, ₹10 crore average daily traded value

The liquidity floors use the mean of close times volume over the last 20 bars and need at least 60 bars of history. These are our settings as of 8 October 2026.

The US list used to be every common stock on NYSE, NYSE American and Nasdaq, about 6,750 symbols taken from Nasdaq Trader’s symbol directories.1 We cut it to the S&P Composite 1500, which combines the S&P 500, MidCap 400 and SmallCap 600 and covers about 90% of US market capitalisation.2 The pauses alone made the case. At 100 symbols a chunk and 7 seconds between chunks, 6,750 symbols spend 67 pauses, almost eight minutes, asleep before any response time is counted. About 1,500 symbols need around 15 pauses, under two minutes. We read the three index lists from Wikipedia, cache them for 24 hours, and fall back to a static S&P 500 list if fewer than 900 symbols come back, so a membership change can reach us late.

Yahoo throttles quietly

In yfinance 1.2.0, the version we run, the bulk download accepts a list of tickers but doesn’t make one request for them. It starts up to twice as many threads as the machine has CPUs and fetches each ticker’s history separately.3 When one of those fetches fails, the library catches the exception, records it and returns an empty frame for that ticker. So a throttled chunk usually raises nothing and comes back empty or full of NaN, with the errors in the log. When the library does raise for throttling, the message is “Too Many Requests. Rate limited. Try after a while.”4

One public report on the yfinance issue tracker describes a job of about 7,000 tickers, close to our old universe, that started getting HTTP 429 responses after roughly 950.5

We treat a chunk as throttled if the exception mentions “rate”, if the result is empty, or if every value in it is NaN. The test is deliberately broad. A chunk of symbols with genuinely no data looks the same and buys a wait it didn’t need, which we accept, because missing a throttle means retrying against a server that has already asked us to stop.

Pacing and backoff on the US side

A threading lock wraps the whole bulk download, so one process runs one sweep at a time. We added it after a US sweep and an India sweep, launched together, tripped the limit and both lost most of their chunks. The lock lives inside the process, so a second process, or a second host behind the same address, isn’t covered.

Each sweep sends 100 symbols per chunk and sleeps 7 seconds between chunks, smaller and slower than our wrapper’s defaults of 250 symbols and 1 second. Each chunk is written to the store as soon as it arrives, so an interrupted sweep keeps what it fetched.

When a chunk looks throttled, we sleep and retry, up to three attempts per chunk. The wait doubles from 45 seconds to a cap of 300 and carries over to the next chunk instead of resetting, so a long throttle isn’t probed again every 45 seconds.

Event Wait before the next request
Chunk A, attempt 1 throttled 45 s
Chunk A, attempt 2 throttled 90 s
Chunk A, attempt 3 throttled Skip chunk A, normal 7 s pause
Chunk B, attempt 1 throttled 180 s
Chunk B, attempt 2 throttled 300 s, the cap
Any chunk that succeeds Reset to 45 s for the next throttle

Skipped symbols are missing from the store, so the next refresh asks for them again. A background thread runs that refresh for both markets every six hours, starting 30 seconds after boot. It uses the same cache-first function and the same lock as a user’s scan, so it only fetches stale or missing stocks and can’t collide with a user’s sweep.

Single-stock lookups go through a separate provider. It waits at least 0.12 seconds between calls, which caps it at about eight a second, caches results for five minutes and remembers failures for 60 seconds. The gap is enforced under a lock that is held through the sleep. The provider has eight worker threads, and if they slept outside the lock, every waiting thread would wake at the same moment and fire together.

What we’d change is jitter. Amazon’s Builders’ Library advises capped exponential backoff with jitter, so that clients that failed together don’t retry together.6 In Marc Brooker’s simulations on the AWS Architecture Blog, backoff without jitter performed worst of the approaches he compared.7 We rarely have clients failing together today, but we will once two hosts share an address. Brooker’s decorrelated variant never waits less than the base, which suits a limit we want to outlast.

import random

def next_wait(previous: float, base: float = 45.0, cap: float = 300.0) -> float:
    """Decorrelated jitter: random, never below the base, never above the cap."""
    return min(cap, random.uniform(base, previous * 3))

One file per trading day on the NSE side

NSE publishes an end-of-day file, the bhavcopy, for every trading day, listing each security’s open, high, low and close prices and traded quantity. One download covers the whole exchange. The full bhavcopy we fetched for this post, from February 2025, was about 310 KB of CSV.

We read two formats. The primary is the full bhavcopy with delivery data, a CSV named sec_bhavdata_full_DDMMYYYY.csv whose headers and values carry stray spaces that we strip. The fallback is the zipped UDiFF common bhavcopy, part of the standardised, ISO-tagged file formats NSE has been moving its members to since late 2022. NSE has retired old file formats before, on announced dates that moved at least once,8 so we keep a parser for each format.

Each parsed day has to contain more than 100 rows, so a truncated download or an error page that happens to parse as CSV counts as no data. It’s only a floor. A file with 101 rows would pass, and we don’t yet compare row counts between days.

Parsed days are cached on disk, one file per day, in Parquet when pyarrow is installed and as a pickle otherwise. A rebuild walks backwards from today through a budget of 385 calendar days (1.5 times the 250-day target, plus 10), skipping weekends without a request and pausing 0.4 seconds after each network fetch. Cached days cost nothing.

When a missing file means a holiday

A 404 for a weekday means the day was a holiday, the file hasn’t been published yet, or NSE moved it. Without a waiting period, an early-evening run would record today as a holiday permanently, so we wait two days.

What happened What we do Why
A file parsed with more than 100 rows Cache the day The data is there
Both formats returned 404, and the date is at least two days old Cache an empty marker, meaning holiday Files appear after the close, so a two-day-old 404 is a real absence
Both formats returned 404, and the date is under two days old Cache nothing and ask again next run The file may not be published yet
A 403, a 5xx, a timeout or a parse failure Cache nothing It says something about the request or the server and nothing about the file
Saturday or Sunday Skip without a request NSE rarely trades at weekends

Two gaps remain. NSE held a full session on Saturday 1 February 2025 for the Union Budget,9 and we fetched that day’s bhavcopy to confirm it exists. Our builder never asks for it, because it skips any date where weekday() >= 5. The next weekend session on NSE’s calendar is Muhurat trading on Sunday 8 November 2026.10 Whether a one-hour festive session belongs in a daily series is a fair question for a calendar to answer.

The other gap is that a renamed file looks exactly like a holiday. If NSE moved both formats tomorrow, every new trading day would return 404 twice and, two days later, be cached as a holiday. The history would stop growing and nothing would alert. The fix is to check each miss against the exchange’s holiday list and raise an alert for a weekday miss that isn’t on it.

Holiday lists we got wrong

While writing this post we checked the one module that does use a holiday list. It answers “is the Indian market open right now?” from a hand-typed set of dates for 2025 and 2026, and we compared its 2026 entries with NSE’s circular for the year, NSE/CMTR/71775.10

Check Result
Dates for 2026 in our list 15
Dates that match NSE’s circular 8
Non-matching dates that were ordinary NSE trading days 6
Non-matching dates that fall on a Saturday 1
NSE weekday holidays our list misses 8 of 15

NSE then added a 16th weekday closure, 15 January 2026, for municipal elections in Maharashtra, with three days’ notice.11

The module only gates the trading loop’s market-open check and a status field, so our price history was unaffected. The list itself marked six trading days as holidays and missed nine closures, counting 15 January, and we’re replacing it with dates loaded from the circulars.

Freshness checks based on calendar days have the same weakness. A rule that a stored stock is fresh if its last bar is no more than four calendar days old, one day plus a three-day pad to cover weekends, is too loose midweek. On a Friday, a file whose last bar is Monday passes, so a scan can run without three sessions of bars, and a refresh that uses the same test leaves the file alone.

Both problems have the same fix. Work out the last completed session from an exchange calendar and compare dates exactly. For NYSE, a maintained library such as exchange_calendars already does this. Its README lists NYSE and the Bombay Stock Exchange and has no NSE calendar,12 so for NSE we’ll load the dates from the circulars, keep them in version control with the circular’s reference, and alert when the files and the list disagree.

Atomic writes, and where we skipped them

The backtest statistics, the scan results, each stock’s file in the US store and each day’s file in the bhavcopy cache are all written the same way. We write the new content to a temporary file in the same directory, then rename it over the target with os.replace, so a reader sees the old file or the new one and never half of each. The sketch below adds the two fsync calls we don’t make yet.

import os
from pathlib import Path

def write_atomically(path: Path, data: bytes) -> None:
    tmp = path.with_name(f"{path.name}.tmp{os.getpid()}")
    with open(tmp, "wb") as f:
        f.write(data)
        f.flush()
        os.fsync(f.fileno())        # not in our code yet
    os.replace(tmp, path)           # atomic on POSIX when it succeeds
    dir_fd = os.open(path.parent, os.O_RDONLY)
    try:
        os.fsync(dir_fd)            # not in our code yet: persists the rename
    finally:
        os.close(dir_fd)

Python’s documentation says that if os.replace succeeds the rename is atomic, which is a POSIX requirement, and that it may fail when the source and target are on different filesystems.13 That’s why the temporary file sits next to the target. The process ID in its name stops two processes saving the same stock from overwriting each other’s half-written temp file, and the last rename wins.

There are two caveats. The first is that atomic isn’t the same as durable. Dan Luu’s survey of file consistency notes that POSIX’s promise about rename applies to normal operation and says nothing about crashes.14 Because we don’t fsync, a power cut at the wrong moment can leave an empty or stale file. Ext4’s delayed allocation showed this in 2009, when files rewritten just before a crash came back empty, and applications were told to call fsync when data has to reach the disk.15

For our caches that risk is acceptable, because every reader treats an unreadable file as missing. The bhavcopy cache deletes a file it can’t read and fetches the day again. The US store skips it, and the next refresh replaces it. A torn file costs one download.

The second caveat is that we skipped the pattern where it seemed not to matter. The cached universe lists use an ordinary write and rely on the reader, which treats a JSON parse error as a cache miss. That works, but it leaves the codebase with two safety models for files, and whoever adds the next cache has to know which one applies.

Splits and bonus issues we don’t adjust for yet

The bhavcopy records prices as they traded that day. It doesn’t adjust earlier prices for corporate actions, and neither do we yet. When a company issues one bonus share for every share held, the price roughly halves overnight, and our series shows a 50% fall in a day.

Everything downstream believes it. RSI collapses, the 20-day and 50-day averages sit far above the price, the Bollinger Bands widen sharply, and detectors that look for oversold bounces or gaps to fill can fire on an event that changed nothing about the company. A backtest that holds through the ex-date books a stop-out that never happened. The detectors look back up to 60 bars, so the distortion lasts about three months of sessions. The price floor can also remove the stock. A ₹30 stock trades near ₹15 after a one-for-one bonus, below our ₹20 minimum. Traded value, price times shares, barely changes with a split or bonus, so that half of the filter is unaffected.

Yahoo covers part of this on the US side. Its price history normally arrives split-adjusted, and yfinance documents a repair option for the cases where Yahoo misses a split, which we don’t turn on.16 The bulk sweep passes auto_adjust=False and keeps only open, high, low, close and volume, so its prices aren’t adjusted for dividends. The single-stock path uses yfinance’s default, auto_adjust=True, which adjusts all four prices.17 The backtest reads the single-stock path and the full-market scan reads the bulk store, so the same stock can carry slightly different histories in each. In India the gap is wider, because the backtest reads adjusted prices from Yahoo while the live scan reads raw bhavcopy prices.

The fix is a table of corporate actions. For each split or bonus, multiply all earlier prices by a factor, 0.5 for a two-for-one split in Yahoo’s description of its own adjusted close,18 and divide earlier volumes by the same factor. We’d store the raw bars and apply the factors when reading, so that correcting the table never means rewriting the cache.

How we checked the figures in this post

The settings, thresholds and file formats come from our code as of 8 October 2026. The pause totals are arithmetic on those settings and count sleep time only. We haven’t benchmarked end-to-end sweep times or measured how often Yahoo throttles us.

For the holiday comparison, we matched the 15 dates for 2026 in our list against NSE circular NSE/CMTR/71775, dated 12 December 2025, and its amendment NSE/CMTR/72260, dated 12 January 2026. We didn’t check the 2025 entries. For the weekend session, we fetched the full bhavcopy for 1 February 2025 from NSE’s archive and confirmed that it holds data dated that day.

What we’d keep, and what we’d change

We’d build these parts the same way again:

  • A process-wide lock, small chunks with real pauses, and backoff that carries across chunks.
  • Empty and all-NaN results treated as throttling.
  • One request per NSE trading day, two parsers and a row-count floor.
  • Holidays cached only on a double 404 at least two days old, and transient errors never cached.
  • A temporary file and a rename for anything read while it may be written, with readers that treat an unreadable file as missing.

And we’d change these:

  • Add jitter, and move the lock out of the process before we run a second one.
  • Replace the weekday check and the four-day pad with exchange calendars, loading NSE’s holidays from its circulars, and alert when files and calendars disagree.
  • Fsync the file and its directory for anything that isn’t a cache.
  • Adjust NSE prices and volumes for splits and bonus issues, with one adjustment policy on every path.
  • Compare each day’s row count with the previous day’s.

  1. Nasdaq Trader, “Symbol Directory Definitions”, https://www.nasdaqtrader.com/trader.aspx?id=symboldirdefs ↩

  2. S&P Dow Jones Indices, “S&P Composite 1500”, https://www.spglobal.com/spdji/en/indices/equity/sp-composite-1500/ ↩

  3. yfinance on GitHub, “multi.py” (the download function), https://github.com/ranaroussi/yfinance/blob/main/yfinance/multi.py ↩

  4. yfinance on GitHub, “exceptions.py”, https://github.com/ranaroussi/yfinance/blob/main/yfinance/exceptions.py ↩

  5. yfinance on GitHub, “Issue 2128”, opened 15 November 2024, https://github.com/ranaroussi/yfinance/issues/2128 ↩

  6. Marc Brooker, “Timeouts, retries, and backoff with jitter”, Amazon Builders’ Library, https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/ ↩

  7. Marc Brooker, “Exponential Backoff And Jitter”, AWS Architecture Blog, 4 March 2015, https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/ ↩

  8. NSE circular NSE/MSD/56202, “Standardisation of file formats for MII-Member Interface, extension of timelines”, 29 March 2023, https://archives.nseindia.com/content/circulars/MSD56202.pdf ↩

  9. NSE circular NSE/CMTR/65729, “Live Trading Session on February 01, 2025, Presentation of Union Budget”, 23 December 2024, https://archives.nseindia.com/content/circulars/CMTR65729.pdf ↩

  10. NSE circular NSE/CMTR/71775, “Trading holidays for the calendar year 2026”, 12 December 2025, https://archives.nseindia.com/content/circulars/CMTR71775.pdf ↩↩

  11. NSE circular NSE/CMTR/72260, “Trading Holiday on January 15, 2026 on account of Municipal Corporation Election in Maharashtra in Capital Market Segment”, 12 January 2026, https://archives.nseindia.com/content/circulars/CMTR72260.pdf ↩

  12. exchange_calendars on GitHub, “README”, https://github.com/gerrymanoim/exchange_calendars ↩

  13. Python documentation, “os.replace”, https://docs.python.org/3/library/os.html#os.replace ↩

  14. Dan Luu, “Files are hard”, https://danluu.com/file-consistency/ ↩

  15. Jonathan Corbet, “ext4 and data loss”, LWN.net, 11 March 2009, https://lwn.net/Articles/322823/ ↩

  16. yfinance documentation, “Price repair”, https://ranaroussi.github.io/yfinance/advanced/price_repair.html ↩

  17. yfinance documentation, “yfinance.download”, https://ranaroussi.github.io/yfinance/reference/api/yfinance.download.html ↩

  18. Yahoo Help, “What is the adjusted close?”, https://help.yahoo.com/kb/SLN28256.html ↩

Frequently asked questions

How do I avoid Yahoo Finance rate limits when using yfinance?

Run one bulk download at a time, keep chunks small with real pauses between them (we use 100 symbols and 7 seconds), and treat empty or all-NaN results as throttling. When throttled, back off for minutes. We start at 45 seconds and cap the wait at 300.

What is the NSE bhavcopy?

It is NSE’s end-of-day file for a trading day, with each traded security’s open, high, low and close prices and traded quantity. One file covers the whole exchange, so a year of history takes about 250 downloads.

How can I tell an NSE holiday from a missing bhavcopy?

We treat a day as a holiday only when both file formats return 404 and the date is at least two days old, and we never cache 403s, server errors or timeouts. A safer design also checks every miss against NSE’s published holiday circular.

Are NSE bhavcopy prices adjusted for splits and bonus issues?

No. The file reports prices as traded, so a split or bonus shows up as a one-day fall. Adjusting needs a table of corporate actions and a factor applied to earlier prices and volumes.

Is os.replace atomic?

Python’s documentation says a successful os.replace is atomic, as POSIX requires, and that it may fail if the source and target are on different filesystems. Atomic isn’t durable, though. Call fsync on the file and its directory if the data must survive a power cut.

Does NSE ever trade on a weekend?

Occasionally. It held a full session on Saturday 1 February 2025 for the Union Budget, and it has scheduled Muhurat trading for Sunday 8 November 2026.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.