2019-07-26 00:55 已编辑湖南大学算法工程师

关注

初学者|不能不会的NLTK

点击上方蓝色字体，关注AI小白入门哟

跟着博主的脚步，每天进步一点点

本文简绍了NLTK的使用方法，这是一个被称为“使用Python进行计算语言学教学和工作的绝佳工具”。

简介

NLTK被称为“使用Python进行计算语言学教学和工作的绝佳工具”。它为50多种语料库和词汇资源（如WordNet）提供了易于使用的界面，还提供了一套用于分类，标记化，词干化，标记，解析和语义推理的文本处理库。接下来然我们一起来实战学习一波~~

官网地址：http://www.nltk.org/

Github地址：https://github.com/nltk/nltk

实战

1.Tokenize

# 安装：pip install nltkimport nltksentence = 'I love natural language processing!'tokens = nltk.word_tokenize(sentence)print(tokens)['I', 'love', 'natural', 'language', 'processing', '!']
import nltk
sentence = 'I love natural language processing!'
tokens = nltk.word_tokenize(sentence)
print(tokens)

['I', 'love', 'natural', 'language', 'processing', '!']

2.词性标注

tagged = nltk.pos_tag(tokens)print(tagged)[('I', 'PRP'), ('love', 'VBP'), ('natural', 'JJ'), ('language', 'NN'), ('processing', 'NN'), ('!', '.')]

[('I', 'PRP'), ('love', 'VBP'), ('natural', 'JJ'), ('language', 'NN'), ('processing', 'NN'), ('!', '.')]

3.命名实体识别

# 下载模型：nltk.download('maxent_ne_chunker')nltk.download('maxent_ne_chunker')[nltk_data] Downloading package maxent_ne_chunker to[nltk_data]     C:\Users\yuquanle\AppData\Roaming\nltk_data...[nltk_data]   Unzipping chunkers\maxent_ne_chunker.zip.Truenltk.download('words')[nltk_data] Downloading package words to[nltk_data]     C:\Users\yuquanle\AppData\Roaming\nltk_data...[nltk_data]   Unzipping corpora\words.zip.Trueentities = nltk.chunk.ne_chunk(tagged)print(entities)(S I/PRP love/VBP natural/JJ language/NN processing/NN !/.)
nltk.download('maxent_ne_chunker')
[nltk_data] Downloading package maxent_ne_chunker to
[nltk_data]     C:\Users\yuquanle\AppData\Roaming\nltk_data...
[nltk_data]   Unzipping chunkers\maxent_ne_chunker.zip.
True

nltk.download('words')
[nltk_data] Downloading package words to
[nltk_data]     C:\Users\yuquanle\AppData\Roaming\nltk_data...
[nltk_data]   Unzipping corpora\words.zip.
True

entities = nltk.chunk.ne_chunk(tagged)
print(entities)

(S I/PRP love/VBP natural/JJ language/NN processing/NN !/.)

4.下载语料库

# 例如：下载brown# 更多语料库：http://www.nltk.org/howto/corpus.htmlnltk.download('brown')[nltk_data] Downloading package brown to[nltk_data]     C:\Users\yuquanle\AppData\Roaming\nltk_data...[nltk_data]   Package brown is already up-to-date!Truefrom nltk.corpus import brownbrown.words()['The', 'Fulton', 'County', 'Grand', 'Jury', 'said', ...]
# 更多语料库：http://www.nltk.org/howto/corpus.html
nltk.download('brown')
[nltk_data] Downloading package brown to
[nltk_data]     C:\Users\yuquanle\AppData\Roaming\nltk_data...
[nltk_data]   Package brown is already up-to-date!
True

from nltk.corpus import brown
brown.words()

['The', 'Fulton', 'County', 'Grand', 'Jury', 'said', ...]

5.度量

# percision：正确率# recall：召回率# f_measurefrom nltk.metrics import precision, recall, f_measurereference = 'DET NN VB DET JJ NN NN IN DET NN'.split()test = 'DET VB VB DET NN NN NN IN DET NN'.split()reference_set = set(reference)test_set = set(test)print("precision:" + str(precision(reference_set, test_set)))print("recall:" + str(recall(reference_set, test_set)))print("f_measure:" + str(f_measure(reference_set,test_set)))precision:1.0recall:0.8f_measure:0.8888888888888888
# recall：召回率
# f_measure
from nltk.metrics import precision, recall, f_measure
reference = 'DET NN VB DET JJ NN NN IN DET NN'.split()
test = 'DET VB VB DET NN NN NN IN DET NN'.split()
reference_set = set(reference)
test_set = set(test)
print("precision:" + str(precision(reference_set, test_set)))
print("recall:" + str(recall(reference_set, test_set)))
print("f_measure:" + str(f_measure(reference_set,
test_set)))

precision:1.0
recall:0.8
f_measure:0.8888888888888888

6.词干提取(Stemmers)

# Porter stemmerfrom nltk.stem.porter import *# 创建词干提取器stemmer = PorterStemmer()plurals = ['caresses', 'flies', 'dies', 'mules', 'denied']singles = [stemmer.stem(plural) for plural in plurals]print(' '.join(singles))caress fli die mule deni#　Snowball stemmerfrom nltk.stem.snowball import SnowballStemmerprint(" ".join(SnowballStemmer.languages))arabic danish dutch english finnish french german hungarian italian norwegian porter portuguese romanian russian spanish swedish# 指定语言stemmer = SnowballStemmer("english")print(stemmer.stem("running"))run
from nltk.stem.porter import *
# 创建词干提取器
stemmer = PorterStemmer()
plurals = ['caresses', 'flies', 'dies', 'mules', 'denied']
singles = [stemmer.stem(plural) for plural in plurals]
print(' '.join(singles))

caress fli die mule deni

#　Snowball stemmer
from nltk.stem.snowball import SnowballStemmer
print(" ".join(SnowballStemmer.languages))
arabic danish dutch english finnish french german hungarian italian norwegian porter portuguese romanian russian spanish swedish
# 指定语言
stemmer = SnowballStemmer("english")
print(stemmer.stem("running"))

run

7.SentiWordNet接口

# 下载sentiwordnet词典import nltknltk.download('sentiwordnet')[nltk_data] Downloading package sentiwordnet to[nltk_data]     C:\Users\yuquanle\AppData\Roaming\nltk_data...[nltk_data]   Unzipping corpora\sentiwordnet.zip.True# SentiSynsets: synsets(同义词集)的情感值from nltk.corpus import sentiwordnet as swnbreakdown = swn.senti_synset('breakdown.n.03')print(breakdown)print(breakdown.pos_score())print(breakdown.neg_score())print(breakdown.obj_score())<breakdown.n.03: PosScore=0.0 NegScore=0.25>0.00.250.75# Lookup(查看)print(list(swn.senti_synsets('slow')))[SentiSynset('decelerate.v.01'), SentiSynset('slow.v.02'), SentiSynset('slow.v.03'), SentiSynset('slow.a.01'), SentiSynset('slow.a.02'), SentiSynset('dense.s.04'), SentiSynset('slow.a.04'), SentiSynset('boring.s.01'), SentiSynset('dull.s.08'), SentiSynset('slowly.r.01'), SentiSynset('behind.r.03')]happy = swn.senti_synsets('happy', 'a')print(list(happy))[SentiSynset('happy.a.01'), SentiSynset('felicitous.s.02'), SentiSynset('glad.s.02'), SentiSynset('happy.s.04')]
import nltk
nltk.download('sentiwordnet')
[nltk_data] Downloading package sentiwordnet to
[nltk_data]     C:\Users\yuquanle\AppData\Roaming\nltk_data...
[nltk_data]   Unzipping corpora\sentiwordnet.zip.
True

# SentiSynsets: synsets(同义词集)的情感值
from nltk.corpus import sentiwordnet as swn
breakdown = swn.senti_synset('breakdown.n.03')
print(breakdown)
print(breakdown.pos_score())
print(breakdown.neg_score())
print(breakdown.obj_score())

<breakdown.n.03: PosScore=0.0 NegScore=0.25>
0.0
0.25
0.75

# Lookup(查看)
print(list(swn.senti_synsets('slow')))
[SentiSynset('decelerate.v.01'), SentiSynset('slow.v.02'), SentiSynset('slow.v.03'), SentiSynset('slow.a.01'), SentiSynset('slow.a.02'), SentiSynset('dense.s.04'), SentiSynset('slow.a.04'), SentiSynset('boring.s.01'), SentiSynset('dull.s.08'), SentiSynset('slowly.r.01'), SentiSynset('behind.r.03')]
happy = swn.senti_synsets('happy', 'a')
print(list(happy))

[SentiSynset('happy.a.01'), SentiSynset('felicitous.s.02'), SentiSynset('glad.s.02'), SentiSynset('happy.s.04')]

更多用法：http://www.nltk.org/howto/index.html

代码已上传：

https://github.com/yuquanle/StudyForNLP/blob/master/NLPtools/NLTKDemo.ipynb

The End

▼往期精彩回顾▼ 新年送福气|您有一份NLP大礼包待领取
自然语言处理中注意力机制综述
达观杯文本智能处理挑战赛冠军解决方案

长按二维码关注
AI小白入门

ID:StudyForAI

学习AI学习ai(爱)

期待与您的相遇~

你点的每个赞，我都认真当成了喜欢

全部评论

推荐最新楼层

05-02 01:12

博世_车辆运动控制系统中国区_数据开发(实习员工)

软件测试 - 商泰汽车 - 一面面经

Boss 投递面试过程：自我介绍对于汽车行业的了解有哪些是否了解测试流程，具体讲解描述一下 HTTP 和 HTTPS 的区别描述一下 UDP 和 TCP 的区别等价类划分法，具体讲解Python 正则表达式当中的 match 和 search 的区别，具体讲解Python 当中的深拷贝和浅拷贝的区别，具体讲解实习经历中，有哪些提升测试覆盖度和效率的实践，具体举例讲解英语口语和读写能力如何反问环节：部门业务（智能座舱，智能驾驶）软件测试与开发

查看10道真题和解析

点赞评论收藏

分享

不愿透露姓名的神秘牛友

04-30 15:30

只有上岸的那一刻是快乐的

出成绩的时候我在高铁站，我不可置信地确认了好几遍，又急忙给家里人打电话，话还没说出口眼泪就先下来了，我在家人群里说了这件事，并且告诉他们先别告诉别人。 之后给导师发邮件，等了几天没有回复也就不挣扎了，因为我在考研期间等待了太多次，等待分配考场、等待出成绩、等待复试时间、等待拟录取通知。 然后我就开始玩，偶尔写论文，期间朋友圈时不时会弹出别人的上岸信息。我觉得自己不知满足，开始后悔当初没有报更好层次的学校。我突然觉得自己的学校拿不出手，又产生了自卑感，尤其是在我爸妈看似无意的向别人炫耀之后，我觉得很烦躁。 我原本计划等录取通知书来了发个朋友圈，现在我什么也不想做了，随便吧……

热血的小松鼠在干饭：比较是偷走幸福的小偷

点赞评论收藏

分享

04-13 18:10

门头沟学院 Java

以后只用抖音

想熬夜的小飞象在秋招：被腾讯挂了后爸妈以为我失联了

点赞评论收藏

分享

03-31 18:02

门头沟学院 Java

也是拒绝过腾讯的人了

白日梦想家_等打包版：不要的哦佛给我

腾讯开奖345人在聊

点赞评论收藏

分享

04-30 13:10

已编辑

东南大学 C++

momenta系统研发实习生

项目拷打非常细，介绍小厂实习做的项目的时候不仅关注技术细节，还有怎么部署，怎么改代码，你写了多少行，怎么实测等等实际问题，直接汗流浃背，写简历包装吹牛的时候没想到这么多。面试到后面语气越来越弱，最后代码题写出来也有错误，后面反问都不好意思问面试表现，还得沉淀

查看1道真题和解析

点赞评论收藏

分享

评论

点赞

收藏

全站热榜

更多

创作者周榜

更多

正在热议

更多

# 实习要如何选择和准备？ #

49597次浏览 782人参与

# 学历or实习经历，哪个更重要 #

90389次浏览 650人参与

# 大疆求职进展汇总 #

474073次浏览 3182人参与

# 摸鱼被leader发现了怎么办 #

46452次浏览 321人参与

# 潍柴工作体验 #

22311次浏览 18人参与

# 你最满意的offer薪资是哪家公司？ #

20606次浏览 120人参与

# 如果可以，你希望哪个公司来捞你 #

69547次浏览 293人参与

# 你觉得通信/硬件有必要实习吗？ #

97088次浏览 893人参与

# Offer比较，求稳定还是求发展 #

44168次浏览 228人参与

# 来聊聊机械薪资天花板是哪家 #

114905次浏览 721人参与

# 硬件兄弟们甩出你的华为奖状 #

97911次浏览 670人参与

# 找工作，行业重要还是岗位重要？ #

22413次浏览 386人参与

# 金融财会交流会 #

103498次浏览 361人参与

# 机械人与华为的爱恨情仇 #

107984次浏览 923人参与

# 24届硬件人与华为的爱恨情仇 #

122572次浏览 962人参与

# 机械人怎么评价今年的华为 #

193005次浏览 1502人参与

# 运营面经 #

103709次浏览 1202人参与

# 外包能不能当跳板？ #

27728次浏览 192人参与

# 实习工作，你找得还顺利吗？ #

397444次浏览 5431人参与

# 国企/银行/研究所公司爆料 #

126368次浏览 742人参与

牛客网
牛客企业服务