CSV files are plain text, and Python reads them with built-in tools or a single import
A CSV file (comma-separated values) is a text file where each line is a row of data and commas separate the columns. Python can read these files three ways: with the built-in open() function and manual parsing, with the csv module that handles commas and special characters automatically, or with pandas if you need to work with the data afterward. Most people use the csv module or pandas because they handle edge cases — like commas inside quoted fields — without extra work.
The method you pick depends on what you do with the data next. If you just need to read it once and print it, the csv module is fastest. If you plan to filter, sort, or analyze the data, pandas is worth the extra line to import it. The built-in open() function works but requires you to write code to handle commas inside quotes, so it is usually not the best choice unless you are working in an environment where you cannot install packages.
Key Takeaways
- The csv module is Python's standard tool for reading CSV files and handles commas inside quoted fields automatically.
- Pandas is faster for large files or when you need to filter, sort, or analyze the data after reading it.
- The open() function works but requires manual code to handle commas and quotes correctly, so it is rarely the best choice.
- All three methods read the file line by line, so they work on files larger than your computer's memory.
- You must know the file path — either the full path or the path relative to where your Python script is saved.
Reading a CSV file with the csv module
The csv module is part of Python's standard library, so you do not need to install anything. Open your Python file and add this at the top:
import csv
Then use this code to read the file:
with open('myfile.csv') as file: reader = csv.reader(file) for row in reader: print(row)
Replace 'myfile.csv' with the actual name of your file. If the file is in a different folder, use the full path: 'C:\Users\YourName\Documents\myfile.csv' on Windows or '/Users/YourName/Documents/myfile.csv' on Mac. The code reads one row at a time and prints it as a list. Each row is a list of strings, so the first row might print as ['Name', 'Age', 'City'].
If your CSV file has a header row (the first row with column names), you can skip it by adding one line:
with open('myfile.csv') as file: reader = csv.reader(file) next(reader) for row in reader: print(row)
The next(reader) line skips the first row. Now your loop starts at the second row.
Reading a CSV file with csv.DictReader
If you want to access columns by name instead of by position, use csv.DictReader instead of csv.reader. This turns each row into a dictionary where the keys are the column names from the header row.
with open('myfile.csv') as file: reader = csv.DictReader(file) for row in reader: print(row['Name'], row['Age'])
This code assumes your CSV has a header row with columns named 'Name' and 'Age'. Each row is now a dictionary, so you can access the value in the 'Name' column with row['Name'] instead of row[0]. This is clearer and less error-prone because you do not have to count which position each column is in.
If your CSV does not have a header row, you can tell DictReader what the column names are:
with open('myfile.csv') as file: reader = csv.DictReader(file, fieldnames=['Name', 'Age', 'City']) for row in reader: print(row['Name'])
Reading a CSV file with pandas
Pandas is a separate package, so you must install it first. Open your terminal or command prompt and type:
pip install pandas
Once installed, use this code:
import pandas as pd df = pd.read_csv('myfile.csv') print(df)
This reads the entire file into memory as a DataFrame, which is a table-like structure. Pandas automatically treats the first row as the header. The output shows all rows and columns in a formatted table. This is much faster than printing rows one at a time if you have thousands of rows.
Pandas is most useful when you need to filter, sort, or analyze the data. For example, to print only rows where the Age column is greater than 30:
import pandas as pd df = pd.read_csv('myfile.csv') filtered = df[df['Age'] > 30] print(filtered)
Pandas also handles missing values, converts columns to the right data type (numbers stay numbers, text stays text), and lets you save the result back to a CSV with df.to_csv('output.csv'). If your file is very large, pandas loads it all into memory at once, which can be slow. For files larger than a few hundred megabytes, the csv module is faster because it reads one row at a time.
Handling common problems when reading CSV files
If you get a "FileNotFoundError", the file path is wrong. Check that the file name is spelled correctly and that the file is actually in the folder you specified. If the file is in the same folder as your Python script, just use the file name: 'myfile.csv'. If it is in a subfolder, use 'subfolder/myfile.csv'. On Windows, you can also use backslashes: 'subfolder\myfile.csv'.
If the data looks wrong — for example, columns are not separated correctly — the file might use a different delimiter. Some CSV files use semicolons or tabs instead of commas. Tell the csv module what delimiter to use:
reader = csv.reader(file, delimiter=';')
Or with pandas:
df = pd.read_csv('myfile.csv', delimiter=';')
If you see strange characters at the start of the data, the file might be encoded in UTF-8 with a BOM (byte order mark). Add the encoding parameter:
with open('myfile.csv', encoding='utf-8-sig') as file:
Or with pandas:
df = pd.read_csv('myfile.csv', encoding='utf-8-sig')
If a column contains numbers but Python treats them as text, pandas can convert them. Use the dtype parameter to specify the data type:
df = pd.read_csv('myfile.csv', dtype={'Age': int})
Choosing between csv, csv.DictReader, and pandas
Use the csv module if you just need to read the file once and do not need to analyze the data. It is fast, requires no installation, and is built into Python. Use csv.DictReader if you want to access columns by name and do not need to do anything complex with the data afterward. Use pandas if you plan to filter, sort, combine, or analyze the data, or if you need to save the result back to a CSV file.
For small files (under 10,000 rows), the speed difference between these methods is not noticeable. For large files, pandas is usually faster because it is written in C and optimized for this work. The csv module is slightly faster than csv.DictReader because it does less work, but the difference is small. Pick the method that makes your code easiest to read and maintain.
Frequently Asked Questions
What if my CSV file has commas inside the data?
The csv module and pandas both handle this automatically. If a field contains a comma, it is surrounded by quotes in the CSV file. The csv module and pandas recognize the quotes and treat the entire quoted section as one field. You do not need to do anything special.
Can I read only certain columns from a CSV file?
With pandas, use the usecols parameter: df = pd.read_csv('myfile.csv', usecols=['Name', 'Age']). With the csv module, read all columns and then use only the ones you need. Pandas is faster if you have a file with many columns and only need a few.
How do I read a CSV file that is on the internet?
Pandas can read directly from a URL: df = pd.read_csv('https://example.com/myfile.csv'). The csv module requires you to download the file first or use the urllib library to fetch it. Pandas is simpler for this task.
What happens if the CSV file is very large?
The csv module reads one row at a time, so it uses very little memory even for huge files. Pandas loads the entire file into memory, which can be slow or fail if the file is larger than your available RAM. For very large files, use the csv module or tell pandas to read the file in chunks with the chunksize parameter.
Can I write data back to a CSV file after reading it?
With pandas, use df.to_csv('output.csv', index=False) to save the DataFrame to a new CSV file. The index=False parameter prevents pandas from adding a row number column. With the csv module, use csv.writer to write rows back to a file.