Skip to content
TaeyoungKim.dev

pandas read_csv UnicodeDecodeError: Check UTF-8 vs. CP949

PythonWritten 3 min readTaeyoungKim
LinkedInX

read_csv() raises UnicodeDecodeError as soon as you open a Korean-language file. The same CSV opens in another program, so pandas may look broken. First ask which character encoding was used to store the file's bytes, and which encoding is decoding them now? UTF-8 and CP949 represent Korean characters with different bytes.

Reproduce CP949 bytes that cannot be decoded as UTF-8

The diagram isolates the encoding choice. It represents the decoding outcome rather than displaying the Korean data itself.

Create an illustrative CSV in memory so the byte encoding is known, unlike an unknown external file. The Unicode escapes below represent Korean header and cell values; ASCII-only text would not demonstrate the difference.

python
from io import BytesIO
import pandas as pd

raw = "\uc774\ub984,\uc0c1\ud0dc\n\uac00,\uc644\ub8cc\n".encode("cp949")

try:
    pd.read_csv(BytesIO(raw))
except UnicodeDecodeError as error:
    print(type(error).__name__)

table = pd.read_csv(BytesIO(raw), encoding="cp949")
print(table.shape)
print(table.columns.tolist() == ["\uc774\ub984", "\uc0c1\ud0dc"])
print(table.iloc[0].tolist() == ["\uac00", "\uc644\ub8cc"])
text
UnicodeDecodeError
(1, 2)
True
True

The first read cannot interpret CP949 bytes with the default UTF-8 decoding path. The second uses the same cp949 encoding that created the bytes and recovers the expected Korean column names and values, as the two checks confirm. A fresh BytesIO(raw) for each read starts at the beginning. When passing a file path, pandas opens the file anew; if you reuse an already open file object, check its read position.

How can you tell whether a real CSV is CP949?

Do not add encoding="cp949" to every file merely because an error shows a byte position. Check the export setting of the source system, the sender's documentation, and other files from the same batch. If you know a test phrase, decode a small sample with candidate encodings and verify that its headers and values are meaningful. Successful decoding with garbled text is still wrong.

Once confirmed, record the encoding in code and operational documentation:

python
table = pd.read_csv("sample.csv", encoding="cp949")
print(table.columns.tolist())
print(table.head(2).to_string(index=False))

Replace sample.csv with the received file. Checking headers and two rows tests more than whether the parser completed. If the file contains personal data, do not copy head() output into shared logs; reproduce with a de-identified test file.

Why be careful with silent replacement or ignored errors?

Ignoring decode failures or replacing characters may keep a pipeline running while changing a name, address, or category code. That can alter join keys and aggregate categories. For important data, surface the failure, confirm the source encoding, and read the original bytes again.

Not every Korean CSV uses CP949. A UTF-8 file may read correctly with the default path. If a batch mixes encodings, keep per-file source and encoding metadata and quarantine unexpected formats rather than guessing silently. Separate a decoding problem from delimiter or speed settings before changing unrelated parser options.

Key takeaways: match stored and decoding encodings

UnicodeDecodeError means the chosen decoder cannot interpret the file's bytes; it does not mean Korean text itself is invalid. Confirm how the source created the file, choose encoding="cp949" when that is the confirmed format, and verify the decoded headers and values afterward.

Author

TaeyoungKim

Connecting technical foundations with implementation, verification, and production decisions.

#pandas#read_csv#CP949#UTF-8#UnicodeDecodeError

Read next