今天在处理python写入csv中文乱码问题,尝试各种办法,百思不得其解,网上找到的最常用的办法就是
1. encoding='utf-8':无效
with open('data4.csv', 'a', encoding='utf-8', newline='') as csvfile:
fieldnames = (['id', 'name', 'age'])
writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
writer.writerow({'id': '10005', 'name': '张伟', 'age': '32'})
2. encoding='utf-8-sig':有效
with open(filename, 'a', newline='', encoding='utf-8-sig') as f: # 中文需要设置成utf-8格式
f_csv = csv.writer(f)
f_csv.writerow(('城市', '日期', '天气', '风力', '温度', '摄氏度')) # 头部信息
f_csv.writerows(data)
3. 先转码再储存:无效
data.decode('utf-8').encode('gbk')
4. 甚至直接以gbk编码写入:无效
f = codecs.open("xx.csv", 'wb', "gbk")
writer = csv.writer(f)
writer.writerow(["客户名称", "行业类型", "客户联系人", "职位", "联系方式", "邮箱","地址"])
到此为止,中文还是乱码,我甚至开始怀疑我的excel软件是不是坏了,。
WPS:想什么呢,我糖糖WPS怎么可能让你随便转换编码,我又不是Notepad++,我是高贵的office!
打开csv源码发现dialect="excel"
反思一下,是不是因为csv太Low了,于是竟然成功了!!!丝毫不敢代码,赶紧贴出来,以防还有新手朋友摆渡摆到翻船。拿走不谢...
writer = csv.DictWriter(csvfile, fieldnames=fieldnames, dialect="excel")
贴上源码:
class DictWriter:
def __init__(self, f, fieldnames, restval="", extrasaction="raise",
dialect="excel", *args, **kwds):
self.fieldnames = fieldnames # list of keys for the dict
self.restval = restval # for writing short dicts
if extrasaction.lower() not in ("raise", "ignore"):
raise ValueError("extrasaction (%s) must be 'raise' or 'ignore'"
% extrasaction)
self.extrasaction = extrasaction
self.writer = writer(f, dialect, *args, **kwds)
def writeheader(self):
header = dict(zip(self.fieldnames, self.fieldnames))
self.writerow(header)
def _dict_to_list(self, rowdict):
if self.extrasaction == "raise":
wrong_fields = rowdict.keys() - self.fieldnames
if wrong_fields:
raise ValueError("dict contains fields not in fieldnames: "
+ ", ".join([repr(x) for x in wrong_fields]))
return (rowdict.get(key, self.restval) for key in self.fieldnames)
def writerow(self, rowdict):
return self.writer.writerow(self._dict_to_list(rowdict))
def writerows(self, rowdicts):
return self.writer.writerows(map(self._dict_to_list, rowdicts))
吃完饭回来发现是python缓存导致,其实encoding='utf-8-sig'是有效,但当时测了很多遍都没生效,缓存误国啊,dialect默认值就是"excel",不需要修改。