①sklearn CountVectorizer\TfidfVectorizer\TfidfTransformer函数详解
"""**Transform a count matrix to a normalized tf or tf-idf representation.**
Tf means term-frequency while tf-idf means term-frequency times inverse
document-frequency. This is a common term weighting scheme in information
retrieval, that has also found good use in document classification.
The goal of using tf-idf instead of the raw frequencies of occurrence of a
token in a given document is to scale down the impact of tokens that occur
very frequently in a given corpus and that are hence empirically less
informative than features that occur in a small fraction of the training
****The formula that is used to compute the tf-idf for a term t of a document d
in a document set is tf-idf(t, d) = tf(t, d) * idf(t), and the idf is
computed as idf(t) = log [ n / df(t) ] + 1 (if ``smooth_idf=False``), where
n is the total number of documents in the document set and df(t) is the
document frequency of t; the document frequency is the number of documents
in the document set that contain the term t. The effect of adding "1" to
the idf in the equation above is that terms with zero idf, i.e., terms
that occur in all documents in a training set, will not be entirely
(Note that the idf formula above differs from the standard textbook
notation that defines the idf as
idf(t) = log [ n / (df(t) + 1) ]).
If ``smooth_idf=True`` (the default), the constant "1" is added to the
numerator and denominator of the idf as if an extra document was seen
containing every term in the collection exactly once, which prevents
zero divisions: idf(t) = log [ (1 + n) / (1 + df(t)) ] + 1.**
Furthermore, the formulas used to compute tf and idf depend
on parameter settings that correspond to the SMART notation used in IR
as follows:
Tf is "n" (natural) by default, "l" (logarithmic) when
Idf is "t" when use_idf is given, "n" (none) otherwise.
Normalization is "c" (cosine) when ``norm='l2'``, "n" (none)
when ``norm=None``.**
Read more in the :ref:`User Guide `.
norm : {'l1', 'l2'}, default='l2'
Each output row will have unit norm, either:
***- 'l2': Sum of squares of vector elements is 1. The cosine
similarity between two vectors is their dot product when l2 norm has
been applied.***
- 'l1': Sum of absolute values of vector elements is 1.
See :func:`preprocessing.normalize`.
use_idf : bool, default=True
Enable inverse-document-frequency reweighting. If False, idf(t) = 1.
smooth_idf : bool, default=True
Smooth idf weights by adding one to document frequencies, as if an
extra document was seen containing every term in the collection
exactly once. Prevents zero divisions.
sublinear_tf : bool, default=False
Apply sublinear tf scaling, i.e. replace tf with 1 + log(tf).
import numpy as np
import pandas as pd
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.feature_extraction.text import TfidfTransformer
tag_list = ['iphone guuci huawei watch huawei',
'huawei watch iphone watch iphone guuci',
'skirt skirt skirt flower',
'watch watch huawei']
vectorizer = CountVectorizer() #将文本中的词语转换为词频矩阵
X = vectorizer.fit_transform(tag_list) #计算个词语出现的次数
transformer = TfidfTransformer()
tfidf = transformer.fit_transform(X) #将词频矩阵X统计成TF-IDF值
(0, 3) 1
(0, 1) 1
(0, 2) 2
(0, 5) 1
(1, 3) 2
(1, 1) 1
(1, 2) 1
(1, 5) 2
(2, 4) 3
(2, 0) 1
(3, 2) 1
(3, 5) 2
[[0. 0.43531168 0.70484465 0.43531168 0. 0.35242232]
[0. 0.34758387 0.2813991 0.69516774 0. 0.5627982 ]
[0.31622777 0. 0. 0. 0.9486833 0. ]
[0. 0. 0.4472136 0. 0. 0.89442719]]
words = np.array(['flower', 'guuci', 'huawei', 'iphone', 'skirt','watch'])
x1=np.array([[0/5, 1/5, 2/5, 1/5, 0/5, 1/5],
[0, 1/6, 1/6, 2/6, 0, 2/6],
[1/4, 0, 0, 0, 3/4, 0],
[0, 0, 1/3, 0, 0, 2/3]])
idfs1=np.array([ 0.6931472,0.2876821,0,0.2876821,0.6931472,0]
0.000000 0.057536 0.0 0.057536 0.00000 0.0
0.000000 0.047947 0.0 0.095894 0.00000 0.0
0.173287 0.000000 0.0 0.000000 0.51986 0.0
0.000000 0.000000 0.0 0.000000 0.00000 0.0
2.上文提到TfidfTransformer函数存在一个参数 smooth_idf ,这主要是为了对得到的IDF权重进行平滑,如果 smooth_idf =true
,则IDF计算公式为idf(t) = log [ (1 + n) / (1 + df(t)) ] + 1,对上面的words我们可以得到IDF矩阵:
idfs1=np.array([1.91629073,1.51082562,1.22314355,1.51082562, 1.91629073, 1.22314355])
0.000000 0.302165 0.489257 0.302165 0.000000 0.244629
0.000000 0.251804 0.203857 0.503609 0.000000 0.407715
0.479073 0.000000 0.000000 0.000000 1.437218 0.000000
0.000000 0.000000 0.407715 0.000000 0.000000 0.815429
3. 如果 smooth_idf =false
,则IDF计算公式为idf(t) = log [ n / df(t) ] + 1,对上面的words我们同样可以得到IDF矩阵:
idfs3=[2.386294, 1.693147, 1.287682, 1.693147,2.386294,1.287682]
0.000000 0.338629 0.515073 0.338629 0.000000 0.257536
0.000000 0.282191 0.214614 0.564382 0.000000 0.429227
0.596573 0.000000 0.000000 0.000000 1.789721 0.000000
0.000000 0.000000 0.429227 0.000000 0.000000 0.858455
③ python中sklearn的pipeline模块: