WebCrawler:用于获取 Coursera、EdX 和 Udacity 数据的爬虫程序

上传者: 42099070 | 上传时间: 2023-01-03 12:10:20 | 文件大小: 804KB | 文件类型: ZIP
网络爬虫 用于获取 Coursera、EdX 和 Udacity 数据的爬虫程序 要求 Python 2.7 具有以下库: 刮痧 要求 JSON 运行爬虫 Coursera 要从 Coursera 收集数据,请运行: python coursera/scrape_coursera.py edX 要从 edX 收集数据,请导航到edx/目录并运行: scrapy crawl edx 优达学城 要从 Udacity 收集数据,请导航到udacity/目录并运行: scrapy crawl udacity

文件下载

资源详情

[{"title":"( 41 个子文件 804KB ) WebCrawler:用于获取 Coursera、EdX 和 Udacity 数据的爬虫程序","children":[{"title":"WebCrawler-master","children":[{"title":"udacity","children":[{"title":"scrapy.cfg <span style='color:#111;'> 256B </span>","children":null,"spread":false},{"title":".DS_Store <span style='color:#111;'> 6.00KB </span>","children":null,"spread":false},{"title":"udacity","children":[{"title":"pipelines.py <span style='color:#111;'> 287B </span>","children":null,"spread":false},{"title":"spiders","children":[{"title":"__init__.pyc <span style='color:#111;'> 168B </span>","children":null,"spread":false},{"title":"udacity_spider.py <span style='color:#111;'> 3.07KB </span>","children":null,"spread":false},{"title":"udacity_spider.pyc <span style='color:#111;'> 3.56KB </span>","children":null,"spread":false},{"title":"__init__.py <span style='color:#111;'> 161B </span>","children":null,"spread":false}],"spread":true},{"title":"__init__.pyc <span style='color:#111;'> 160B </span>","children":null,"spread":false},{"title":"data","children":[{"title":"items.json <span style='color:#111;'> 736.59KB </span>","children":null,"spread":false},{"title":".DS_Store <span style='color:#111;'> 6.00KB </span>","children":null,"spread":false},{"title":"udacity_3.json <span style='color:#111;'> 61.37KB </span>","children":null,"spread":false},{"title":"udacity.json <span style='color:#111;'> 152.50KB </span>","children":null,"spread":false}],"spread":true},{"title":".DS_Store <span style='color:#111;'> 15.00KB </span>","children":null,"spread":false},{"title":"provider <span style='color:#111;'> 324.25KB </span>","children":null,"spread":false},{"title":"items.py <span style='color:#111;'> 286B </span>","children":null,"spread":false},{"title":"__init__.py <span style='color:#111;'> 0B </span>","children":null,"spread":false},{"title":"item.json <span style='color:#111;'> 1.28KB </span>","children":null,"spread":false},{"title":"settings.py <span style='color:#111;'> 488B </span>","children":null,"spread":false},{"title":"settings.pyc <span style='color:#111;'> 270B </span>","children":null,"spread":false}],"spread":false}],"spread":true},{"title":"edx","children":[{"title":"scrapy.cfg <span style='color:#111;'> 248B </span>","children":null,"spread":false},{"title":".DS_Store <span style='color:#111;'> 6.00KB </span>","children":null,"spread":false},{"title":"edx","children":[{"title":"pipelines.py <span style='color:#111;'> 257B </span>","children":null,"spread":false},{"title":"spiders","children":[{"title":"edx_spider.py <span style='color:#111;'> 2.79KB </span>","children":null,"spread":false},{"title":"__init__.pyc <span style='color:#111;'> 160B </span>","children":null,"spread":false},{"title":".DS_Store <span style='color:#111;'> 6.00KB </span>","children":null,"spread":false},{"title":"__init__.py <span style='color:#111;'> 161B </span>","children":null,"spread":false},{"title":"edx_spider.pyc <span style='color:#111;'> 3.55KB </span>","children":null,"spread":false}],"spread":true},{"title":"__init__.pyc <span style='color:#111;'> 152B </span>","children":null,"spread":false},{"title":"data","children":[{"title":"category_map <span style='color:#111;'> 420B </span>","children":null,"spread":false},{"title":"items.json <span style='color:#111;'> 506.86KB </span>","children":null,"spread":false},{"title":"test.json <span style='color:#111;'> 924.03KB </span>","children":null,"spread":false},{"title":"edx.json <span style='color:#111;'> 311.50KB </span>","children":null,"spread":false}],"spread":true},{"title":"test.json <span style='color:#111;'> 6B </span>","children":null,"spread":false},{"title":".DS_Store <span style='color:#111;'> 15.00KB </span>","children":null,"spread":false},{"title":"items.py <span style='color:#111;'> 421B </span>","children":null,"spread":false},{"title":"__init__.py <span style='color:#111;'> 0B </span>","children":null,"spread":false},{"title":"settings.py <span style='color:#111;'> 443B </span>","children":null,"spread":false},{"title":"settings.pyc <span style='color:#111;'> 254B </span>","children":null,"spread":false}],"spread":true}],"spread":true},{"title":"README.md <span style='color:#111;'> 498B </span>","children":null,"spread":false},{"title":"coursera","children":[{"title":"coursera_requests.py <span style='color:#111;'> 3.07KB </span>","children":null,"spread":false},{"title":"scrape_coursera.py <span style='color:#111;'> 246B </span>","children":null,"spread":false}],"spread":true}],"spread":true}],"spread":true}]

评论信息

免责申明

【只为小站】的资源来自网友分享,仅供学习研究,请务必在下载后24小时内给予删除,不得用于其他任何用途,否则后果自负。基于互联网的特殊性,【只为小站】 无法对用户传输的作品、信息、内容的权属或合法性、合规性、真实性、科学性、完整权、有效性等进行实质审查;无论 【只为小站】 经营者是否已进行审查,用户均应自行承担因其传输的作品、信息、内容而可能或已经产生的侵权或权属纠纷等法律责任。
本站所有资源不代表本站的观点或立场,基于网友分享,根据中国法律《信息网络传播权保护条例》第二十二条之规定,若资源存在侵权或相关问题请联系本站客服人员,zhiweidada#qq.com,请把#换成@,本站将给予最大的支持与配合,做到及时反馈和处理。关于更多版权及免责申明参见 版权及免责申明