02:18
<Dashiva>
Extensible ponies are the best
05:17
<StreVat>
Hello all
05:18
<StreVat>
Just getting started developing a blog here and I wanna take the opportunity to do what I can in html5
05:18
<StreVat>
if anyone has any prized examples, let me know!
05:38
<othermaciej>
if the entire main content of a document is an article, is it appropriate to use <article>?
05:50
<Hixie>
othermaciej: yes
07:43
<boblet>
StreVat: re HTML5 blog, two overviews of HTML5 sectioning elements (of most interest to authors) are:
07:43
<boblet>
http://edward.oconnor.cx/2009/09/using-the-html5-sectioning-elements
07:43
<boblet>
http://boblet.tumblr.com/post/141239118/html5-structure4
07:44
<boblet>
Also, http://html5doctor.com/ has a lot of good articles
07:44
<boblet>
any questions ask here, and someone will generally help you out
07:48
hsivonen
wonders how to make an eclipse project notice its Team function should now belong to MercurialEclipse--not Subclipse
08:45
<Hixie>
thanks to Ms2ger, HTML5 now has an index of elements
08:46
<nessy>
w00t!
08:48
<Philip`>
Rather than index, you should call it a periodic table
08:49
<othermaciej>
would anyone like to help me proofread http://dev.w3.org/html5/decision-policy/decision-policy.html before I send it to the HTML WG mailing list?
08:53
<Philip`>
othermaciej: Shouldn't the bug be closed in 5.b.?
08:54
<othermaciej>
Let me see
08:54
<Philip`>
(It seems bad to leave open bugs hanging around forever)
08:54
<Philip`>
s/in 5.b./somewhere in the path leading up to 5.b./
08:55
<othermaciej>
Philip`: that's more substantive feedback than proofreading - right now by design we don't close it in that case, since the originator could come back at some later point
08:55
<othermaciej>
Philip`: probably good feedback to post on the mailing list once I post this
08:55
<othermaciej>
Philip`: we did discuss a possibility like tagging with a keyword or marking CLOSED with a keyword after some period of time and no reply
08:55
<othermaciej>
PS I just fixed the broken anchor links
08:56
Philip`
reads the bit later on
08:56
<nessy>
othermaciej: that looks very helpful! will read later! gotta go talk about video a11y...
08:56
<Philip`>
Oh, it gets set to RESOLVED, so maybe that's enough
08:56
<othermaciej>
nessy: good luck
08:56
<nessy>
thanks :)
08:56
<othermaciej>
nessy: it will be in VERIFY by that point
08:56
<nessy>
just a 3 min webjam talk :)
08:56
<othermaciej>
er
08:56
<othermaciej>
meant that for Philip`
08:56
<nessy>
no worries :)
08:58
<Philip`>
othermaciej: "its usually not a good idea to repeatedly reopen the same bug" - s//'/
08:58
<othermaciej>
hmm, is the state really VERIFIED or is it VERIFY? Can't remember
08:59
<zcorpan>
the former
08:59
<othermaciej>
all righty
08:59
<othermaciej>
Philip`: thanks, fixed in my copy
09:00
<hsivonen>
othermaciej: the arrow from 6. WG Decision to 7.a. is not properly anchored
09:00
<Philip`>
othermaciej: "Rasied issue" (in diagram)
09:00
<Hixie>
othermaciej: in step 2, is "A rationale for the change or lack of change (at least enough for the Disposition of Comments).", is saying nothing and just marking the bug fixed acceptable when the bug report is, e.g., a typo?
09:01
<Hixie>
othermaciej: or would the editor have to say "I corrected this because I agree that that is a spelling mistake" or something asinine like that
09:01
<hsivonen>
othermaciej: why does a bug have to include at least one fix suggestion?
09:01
<Philip`>
othermaciej: "settled amicably, If spec" - s/,/./
09:01
<othermaciej>
Hixie: for trivial typos where you accept the comment, no rationale is needed; would you like me to clarify explicitly?
09:01
<othermaciej>
hsivonen: it doesn't have to - that's a suggestion
09:01
<Hixie>
othermaciej: nah, just curious
09:02
<hsivonen>
othermaciej: also, sample spec text needs to be rewritten anyway to wash away copyright
09:03
<othermaciej>
hsivonen: I can remove that part of the suggestion, but in some cases, short pieces of sample spec text have been useful to Hixie
09:03
<othermaciej>
e.g. the XPath issue
09:03
<hsivonen>
othermaciej: oh, OK. in general Hixie has discouraged spec text
09:03
<Hixie>
the XPath issue was a very rare exception
09:03
<othermaciej>
I can imagine that in many cases they'll be useful for Manu's HTML+RDFa draft as well
09:03
<hsivonen>
othermaciej: ok
09:03
<othermaciej>
this policy is meant to apply to all deliverables that reach at least Working Draft status
09:04
<Hixie>
(with the XPath issue, what happened is that the commentor disagreed on an editorial basis, and I couldn't work out why without seeing what they thought was the right answer.)
09:05
<hsivonen>
othermaciej: is Bugzilla flexible enough to let bug reporters change state to CLOSED?
09:06
<hsivonen>
othermaciej: also, this doesn't cover anonymous comments from the WHATWG system
09:06
<othermaciej>
hsivonen: if you make an anonymous comment, then it is your choice to either identify yourself in the bug or let it stay in VERIFY whatever the resolution
09:07
<hsivonen>
othermaciej: I think you need a cutoff timeout for 5.b. Otherwise, there will be huge permathreads on when we can stop waiting.
09:07
<othermaciej>
hsivonen: I am sure MikeSmith can do the necessary bugzilla magic, but if not we can tweak that step
09:07
<othermaciej>
hsivonen: all right - I recommend giving that feedback on the list after I post this
09:07
<othermaciej>
(don't want to make substantial changes right now)
09:07
<othermaciej>
Philip` made the same point and I agree
09:08
<hsivonen>
If I make that point on the list, there's a risk of a huge permathread right there
09:09
<othermaciej>
I'm inclined to say you get a month to reply to the editor's disposition but I don't want to change it right now since my co-chairs signed off on the substance of what is there now
09:10
<Hixie>
what's 5b?
09:10
<Hixie>
isn't that the "nobody replied" state?
09:10
<othermaciej>
yes
09:11
<Hixie>
there's no point having a cut-off for that -- someone not replying until after the cut-off is the same as someone filing a new bug
09:11
<Hixie>
in fact in general there's no difference between someone disagreeing and someone filing a new bug
09:11
<Hixie>
so there's no point having a cut-off
09:13
<hsivonen>
in that case, there should be some mechanism for deflecting dupes
09:14
<Dashiva>
Is there a consistent meaning to squares, diamonds and ovals?
09:14
<othermaciej>
dupes should be resolved DUPLICATE unless they contain new information
09:15
<othermaciej>
rounded rects are start or end points, rectangles are process steps, diamonds are decisions
09:15
<othermaciej>
per standard flowchart usage
09:15
Philip`
notes that the change proposal deadline doesn't particularly important, because you could always raise an identical issue months later and attach your change proposal
09:15
<Philip`>
s/doesn't/doesn't seem/
09:15
<Dashiva>
Escalation step 2b doesn't seem like a decision
09:16
<othermaciej>
oh that's because I forgot to add an arrow pointing to 2a
09:16
<othermaciej>
fixing...
09:18
<hsivonen>
What happens if the person who escalated a bug to ISSUE is willing to settle amicably but someone else wants to stir up trouble? does the other person have to restart the whole process?
09:19
<hsivonen>
there seems to be a loophole in the case where the editor has a co-conspirator:
09:20
<othermaciej>
then the issue stays RAISED and someone has to make a Change Proposal (presumably not the person who first raised it)
09:20
<hsivonen>
the co-conspirator escalates an issue, fails to produce a proposal and the issue gets closed and deferred to the next version of HTML
09:20
<othermaciej>
amicable resolution does not apply if somebody objects, but then it's up to them to drive the issue
09:20
<Dashiva>
Maybe there should be an option for bug reporters to close as INVALID if they realize they misunderstood or have been convinced in other channels?
09:21
<Dashiva>
Similar to amicable resolution in escalation
09:21
<othermaciej>
I'm assuming that is allowed, but I guess it could be spelled out more explicitly
09:22
<othermaciej>
hsivonen: what's the alternative - that someone else would have escalated it later when they had more time to write a Change Proposal?
09:22
<hsivonen>
othermaciej: right
09:22
<othermaciej>
hsivonen: multiple people can volunteer to make a Change Proposal, so to do that someone would have to lie about their intent to make one and convince others that it's one they will like
09:23
<hsivonen>
othermaciej: true
09:23
<Dashiva>
Well, with the right timing, maybe the people who care are all busy with other proposals, or on vacation, or something :)
09:24
<hsivonen>
othermaciej: shouldn't Proposal Details be required to come with a copyright waiver?
09:25
<hsivonen>
or a super-permissive license, rather
09:25
<othermaciej>
I'm going to assume no one is going to escalate in bad faith - if we caught someone doing that, we'd likely treat it as a de facto amicable resolution (leaving someone else free to re-raise the issue or object or whatever)
09:25
<othermaciej>
we'd probably also have a stern talk with their AC rep
09:25
<Hixie>
what if they _are_ the AC rep? :-)
09:26
<Hixie>
(i don't understand the issue, personally. Why can't someone just write a proposal later, when they feel like it?)
09:26
<Dashiva>
What if they're an IE?
09:26
<othermaciej>
sure, anyone can write a proposal later - it can still be considered, it just won't be a blocker to wait for it
09:27
<hsivonen>
othermaciej: oh does the decision to defer to the next version of HTML get reversed if someone else re-escaletes the same issue?
09:28
<othermaciej>
hsivonen: I guess it's not very crisply defined as written
09:28
<Hixie>
there's no way to prevent that, as far as i can see; someone can always claim that the issue they are raising is subtly different
09:28
<Hixie>
(which is fine)
09:28
<Hixie>
ok, index is checked in. File bugs if you find any. I'm going to bed. nn.
09:29
<hsivonen>
ok. in that case, the process isn't vulnerable to trying to pre-empt decision by getting them deferred but the process is vulnerable to Denial of Productivity escalation flood
09:29
<othermaciej>
Dashiva: if you don't mind saying, what's your real world name (so I can credit you for finding a mistake in my commit message)
09:29
<hsivonen>
which may be something that isn't worth fixing unless a flood happens
09:29
<Dashiva>
othermaciej: Magnus Kristiansen
09:30
<othermaciej>
hsivonen: the intent is that for your escalation flood to really have an effect, you have to do some work - someone re-raising the same timed-out issue over and over would be bad-faith behavior though, which we'd have to address if someone did it
09:30
<hsivonen>
othermaciej: OK.
09:30
<othermaciej>
I think I'm going to mail this to public-html, if there are no more obvious typos
09:36
<Philip`>
othermaciej: I think there's a big typo - you accidentally wrote whole pages of text, instead of "Hixie decides everything."
09:36
<othermaciej>
lol
09:37
<Dashiva>
So how does the process ensure that someone defined D.E. before a WG decision is made?
09:38
<hsivonen>
Dashiva: excellent point
09:39
<zcorpan>
D.E. == Dead End?
09:39
<hsivonen>
hehe
09:43
<zcorpan>
othermaciej: you have both "Full Working Group" and "full Working Group"
09:44
<othermaciej>
the Full shouldn't be capitalized
09:44
<othermaciej>
thanks
10:19
<gsnedders|work>
Hmm, pne of the issues on Web DOM Core says, "should we remove error checking altogether?", which looks to be impossible (as that will break websites!)
10:20
<gsnedders|work>
I wonder if it's more possible to go the other way, and require the DOM to represent well-formed XML
10:20
<gsnedders|work>
The current middle-ground is just idiotic
10:20
<hsivonen>
gsnedders|work: what site patterns rely on DOM Core error checking?
10:20
<jgraham>
gsnedders|work: I doubt it
10:20
<hsivonen>
gsnedders|work: you can't require a DOM to be serializable without breaking L1 colon usage
10:21
<gsnedders|work>
hsivonen: A lot of those that use e.g., document.createElement("<div class=foobar>") for IE rely upon that throwing an exception to have their code work in other browsers
10:21
<gsnedders|work>
hsivonen: Ah yeah, that's true
10:21
<gsnedders|work>
hsivonen: And would break RDFa, totally :P
10:21
<othermaciej>
gsnedders|work: wait, that works in IE?
10:21
<gsnedders|work>
othermaciej: Yes.
10:22
othermaciej
facepalms
10:22
<jgraham>
gsnedders|work: But there are downsides too, rihgt? ;)
10:22
<gsnedders|work>
othermaciej: Firefox supports it without attributes in quirks mode
10:22
<hsivonen>
gsnedders|work: ouch. so we can't make that syntax work in other browsers?
10:22
<hsivonen>
gsnedders|work: IIRC, Gecko had some half-way magic on that point
10:22
<othermaciej>
we've never supported anything like that in WebKit afaik
10:22
<gsnedders|work>
hsivonen: We could, but I don't like the idea of making DOM rely upon having an HTML tokenizer around
10:22
<hsivonen>
gsnedders|work: does the Web expect document.createElement("<div>") to create a div element, though?
10:22
<gsnedders|work>
othermaciej: Nor does Opera. It doesn't cause any site compat issues I've found
10:22
<gsnedders|work>
hsivonen: No
10:23
<hsivonen>
hmm.
10:23
<gsnedders|work>
Anyhow, lunch
10:23
<jgraham>
hsivonen: (except in IE)
10:23
<othermaciej>
if sites depended on that working and not throwing, I would expect compat bugs, since scripts unexpectedly throwing tends to cause severe failure modes
10:25
<hsivonen>
jgraham, gsnedders|work: my recollection was right. Gecko does have the half-way magic in the quirks mode
10:25
<hsivonen>
http://mxr.mozilla.org/mozilla-central/source/content/html/document/src/nsHTMLDocument.cpp#1208
10:26
<hsivonen>
if the argument starts with < and ends with >, those characters are stripped in the quirks mode before proceeding
10:27
<othermaciej>
indeed I see the code doing that
10:28
<hsivonen>
(I mentioned what the code does in case people are prohibited from looking at code on MXR.)
10:30
<hsivonen>
https://bugzilla.mozilla.org/show_bug.cgi?id=245274
10:30
<hsivonen>
https://bugzilla.mozilla.org/show_bug.cgi?id=245274#c7
10:32
<jgraham>
Interesting
10:32
<jgraham>
Is there other DOM stuff that depends on the quirkiness of the document?
10:33
<hsivonen>
jgraham: IIRC, yes
10:33
<othermaciej>
interesting - we never got those reports
10:33
<othermaciej>
(though we have had many bug reports from IBM on problems with their enterprise web apps)
10:34
<jgraham>
(I note in his absence that gsnedders|work found some stuff that used this to detect IE vs non IE and give them diffeent code paths)
10:34
<jgraham>
(and so broke in Gecko)
10:34
<jgraham>
(but I'm not sure how much)
10:34
<othermaciej>
wild
10:35
<hsivonen>
getElementsByClassName()
10:36
<hsivonen>
scroll info for body
10:37
<hsivonen>
document.all
10:37
<hsivonen>
color parsing
10:37
<hsivonen>
in attribute values
10:39
<hsivonen>
various attributes in tables
10:41
<zcorpan>
http://forums.whatwg.org/viewtopic.php?p=5274#5281
10:42
<zcorpan>
othermaciej: "These components are not absolutely mandatory for a bug reports." s/a bug/all bug/
10:43
<othermaciej>
zcorpan: thanks
10:43
<zcorpan>
othermaciej: i thought bugs with insufficient information would be RESOLVED NEEDSINFO?
10:43
<othermaciej>
zcorpan: I guess it depends on the nature of what is missing
10:44
<othermaciej>
zcorpan: can you send that feedback to public-html please? I'd rather have substantive feedback on the list
10:45
<zcorpan>
ok
11:08
<gsnedders|work>
othermaciej: We've never had bug reports about it either
11:08
<gsnedders|work>
As far as I can tell, Gecko's half-way state is the worst for web compat
11:13
<jgraham>
gsnedders|work: But possibly the best for intranet compat :(
11:22
<gsnedders|work>
hsivonen: How many of the quirks do you think need to be in Web DOM Core and not HTML 5?
11:23
<othermaciej>
anything specific to HTML APIs should clearly be in HTML5
11:23
<othermaciej>
core DOM quirks, I could see an argument either way
11:23
<othermaciej>
though it would be nice if DOM Core itself could be kept free of quirks-mode-specific behaviors
11:25
<gsnedders|work>
That's what I think too. I'm just trying to think of whether there are any quirks that need to be in DOM Core
11:27
<gsnedders|work>
The fact that neither WebKit nor Opera supports the <…> createElement fun is encouraging, though it's unclear whether intranet stuff relies upon it stil.
11:27
<hsivonen>
gsnedders|work: they can all be in HTML5 if you accept the delta spec section in HTML5 that already modifies DOM Core
11:27
<othermaciej>
well that createElement quirk would be a candidate, but I'm hoping it's safe to just drop
11:27
<othermaciej>
the other things hsivonen mentioned are specific to HTML APIs I think not DOM Core APIs
11:27
hsivonen
wonders how much IBM intranet stuff relies on IE Custom Tags
11:28
<hsivonen>
depends on whether one views document.all as HTML or Core
11:28
<othermaciej>
true; you could put document.all in Core
11:28
<hsivonen>
though one might argue that Gecko is wrong to make document.all conditional on the quirkiness
11:28
<othermaciej>
in WebKit I believe we expose undetectable document.all in all modes
11:29
<gsnedders|work>
hsivonen: There's no way around having HTML 5 modify certain behaviours for HTMLDocumnent and HTMLElement AFAIK. I'd rather document.all was HTMLDocument.all
11:29
<gsnedders|work>
I really don't see any reason for that to create onto Document.all
11:29
<hsivonen>
gsnedders|work: how does that help when HTML5 requires all document objects to implement HTMLDocument and SVGDocument anyway?
11:30
<othermaciej>
since HTML5 requires all document objects to implement HTMLDocument and SVGDocument, there's not much practical benefit to moving APIs from HTML5 to DOM Core
11:30
<hsivonen>
keeping Core clean and applying a delta in HTML5 reminds me of the latest TAG minutes
11:30
<othermaciej>
just a potential question for spec purity
11:30
<othermaciej>
for cases where HTML5 modifies the behavior of existing DOM Core methods, that should ideally be in DOM Core
11:31
<othermaciej>
those aren't quirks, just changes (applicable only to HTML documents)
11:31
<gsnedders|work>
othermaciej: Some of those are specific to HTML documents though
11:31
<gsnedders|work>
othermaciej: And it seems silly for those to be in DOM Core when they only apply to HTML documents
11:31
<othermaciej>
I wonder if it would be useful to expose the HTMLness bit of the Document in some more explicit way
11:32
<othermaciej>
gsnedders|work: at the very least DOM Core should be written in a way that HTML5 doesn't have to *contradict* it
11:32
<othermaciej>
I think it would be fine to have an extension point instead of specifying html document behaviors directly
11:32
<othermaciej>
though that creates a bit more risk that things fall through the cracks between the documents
11:33
<gsnedders|work>
othermaciej: How can we avoid having HTMLDocument override things like createElement though, to do case-insensitive creation?
11:33
<othermaciej>
I think it would be fine for DOM Core to be aware of the idea of XML documents and HTML documents as different things
11:33
<othermaciej>
gsnedders|work: that's a property of being an "html document" (as opposed to an xml document), not a property of the HTMLDocument interface
11:54
<hsivonen>
based on the test case Opera and WebKit don't do the magic
11:54
<zcorpan>
the test might be bogus for opera
11:54
<othermaciej>
we do look at the namespaceURI parameter to createDocument() to decide what document interface to create
11:55
<othermaciej>
we do not yet do the HTML5 thing of all interfaces on all documents
11:56
hsivonen
goes file a bug
12:00
<zcorpan>
http://software.hixie.ch/utilities/js/live-dom-viewer/saved/263 has a correct test for opera
12:01
<zcorpan>
when i think about it, i'm not sure opera has the htmlness bit on the document, but has it on elements instead
12:02
<zcorpan>
uh, that test is still bogus
12:03
<zcorpan>
a HEAD element is inserted...
12:03
zcorpan
gives up
12:03
<gsnedders|work>
Honestly, can't you write a non-bogus test?
12:03
<zcorpan>
no
12:05
<hsivonen>
ttps://bugzilla.mozilla.org/show_bug.cgi?id=520969
12:05
<hsivonen>
https://bugzilla.mozilla.org/show_bug.cgi?id=520969
12:07
hsivonen
wishes Congress took the Eastern District of Texas out of business already
12:08
<hsivonen>
or SCOTUS if the Congress can't get their act together
12:08
<zcorpan>
http://software.hixie.ch/utilities/js/live-dom-viewer/saved/264 is non-bogus for opera (but bogus for others)
13:17
Philip`
wonders why http://www.quirksmode.org/webkit.html includes Konqueror 3.5 in its list of WebKits
14:03
<zcorpan>
http://simon.html5.org/dump/html+js+css+atom.html - a file that is interpreted as html, as js, as css and as atom at the same time
14:03
<zcorpan>
works in firefox at least
14:03
<zcorpan>
the atom part doesn't seem to work in opera
14:05
<Rik|work>
Philip`: I even wonder if Konqueror is on a phone
14:05
<TabAtkins>
Heh, interesting zcorpan.
14:06
<zcorpan>
hmm i have a css syntax error at the end
14:07
<zcorpan>
solved
14:07
<jgraham>
You need a hobby
14:07
<jgraham>
A better one I mean
14:07
<Philip`>
zcorpan: Make it a PNG too
14:08
<zcorpan>
Philip`: how?
14:08
<Philip`>
That's your problem, not mine :-p
14:08
<zcorpan>
i tried making it SVG too but couldn't find a way to link it in and being rendered as SVG
16:53
<mpilgrim>
i committed a few fixes to the python3 branch of html5lib
16:53
<mpilgrim>
it crashes less now
16:54
<mpilgrim>
i haven't tried anything crazy like running the test suite
16:56
<jgraham>
mpilgrim: I noticed, thanks
16:56
<mpilgrim>
what is the status of the python3 port?
16:57
<jgraham>
It was an experimental 1 weekend hack to see if it would work
16:57
<mpilgrim>
that's what i suspected
16:57
<jgraham>
It is not up to date wrt to the trunk
16:57
<mpilgrim>
how far behind is it?
16:58
<jgraham>
More to the point I don't really know how to do parallel maintainance of the python 3 and python 2 versions without lots of manual work porting patches
16:58
<mpilgrim>
neither do i
16:58
<jgraham>
mpilgrim: A few months but not a huge amount has happened in that time
16:58
<gsnedders>
s/in/to the Python port in/
16:59
<mpilgrim>
i managed to get it to work with a string input
16:59
<mpilgrim>
i can't get it to work with a file
16:59
<mpilgrim>
er, file object
16:59
<jgraham>
mpilgrim: I remember that I couldn't decide how to deal with file input
16:59
<mpilgrim>
it complains that binary files are not supported
16:59
<mpilgrim>
even though i opened the file in non-binary mode
16:59
<jgraham>
hmm
17:00
<jgraham>
The problem I remember is that html5lib really really wants bytestreams
17:00
<mpilgrim>
but if i read the file myself and pass a python string to the parser, it works
17:00
<jgraham>
not files with encodings
17:00
<jgraham>
a string not a byte (whatever the name is)?
17:00
<Philip`>
Is it possible for core parts of html5lib to be written in a common subset of Pythons 2 and 3, and just use version-specific ports for the stuff around the edges?
17:00
<Philip`>
or are the differences more fundamental?
17:01
<jgraham>
Philip`: I think the differences are more non-trivial than that
17:01
<jgraham>
For example .iteritems v .items
17:01
<mpilgrim>
it might be possible to structure some of the code in such a way that it can be automatically converted to (or from) python 3
17:02
<mpilgrim>
that will cover things like syntax differences in importing relative paths within the codebase
17:02
<mpilgrim>
and print statements
17:02
<mpilgrim>
and such
17:02
<mpilgrim>
the big problem is python 2 has "strings" and "unicode strings"
17:02
<mpilgrim>
while python 3 has "bytes" and "unicode strings"
17:03
<mpilgrim>
python 2 implicitly converts between them
17:03
<mpilgrim>
python 3 does not
17:03
<jgraham>
Yeah, so like I said html5lib really wants bytes, always
17:04
<jgraham>
because it really wants to figure out the encoding on its own
17:04
<jgraham>
But then internally everything should be unicode strings
17:05
<gsnedders>
It should cope fine with getting a Unicode string though
17:06
<gsnedders>
And should just do no decoding of any sort whatsoever
17:06
<jgraham>
gsnedders: Well you can't really do that and implement the spec
17:06
<jgraham>
Or maybe you can and treat it like a HTTP header
17:06
<gsnedders>
jgraham: If you already know the encoding you don't need to detect it
17:07
<mpilgrim>
well, some parts of html5lib can certainly be made python-version-independent
17:07
<jgraham>
Yeah fair enough
17:07
<jgraham>
(that was aimed at gsnedders)
17:07
<gsnedders>
jgraham: It makes no sense to do anything else with Unicode strings
17:07
<mpilgrim>
or at least automatically-convertible
17:07
<mpilgrim>
constants.py, for example, could be automatically converted
17:08
<jgraham>
mpilgrim: I think it should be possible to make most of the code except some parts of inputstream.py automatically convertable
17:08
<jgraham>
But that is just conjecture
17:08
<gsnedders>
It's being transported within memory and that has a defined encoding, so step 1 of the encoding sniffing algorithm applies.
17:09
<jgraham>
gsnedders: (I seem to remember that there was still some reason you wanted a byte array but I don't recall what it was... maybe to do with replacement characters or something?)
17:09
<gsnedders>
jgraham: That was in PHP, no?
17:09
<jgraham>
gsnedders: No
17:25
<mpilgrim>
html5lib/utils.py can be safely auto-converted
17:26
<mpilgrim>
grr
17:26
<mpilgrim>
html5lib/constants.py contains some strings that are later used in %-replacement string formatting
17:28
<zcorpan_>
MikeSmith: did you hear from someone about trouble with bugzilla account?
17:29
<mpilgrim>
html5lib/ihatexml.py can be safely auto-converted
17:31
<jgraham>
mpilgrim: Do the strings correspond to the error messages? I'm not sure that feature provides much value
17:31
<mpilgrim>
python3/html5lib/sanitizer.py appears to be out of date, but i think it can be safely auto-converted too
17:31
<mpilgrim>
yes
17:31
<mpilgrim>
there are strings like
17:31
<mpilgrim>
"cant-convert-numeric-entity":
17:31
<mpilgrim>
_(u"Numeric entity couldn't be converted to character "
17:31
<mpilgrim>
u"(codepoint U+%(charAsInt)08x)."),
17:32
<mpilgrim>
and others like
17:32
<mpilgrim>
"expected-closing-tag-but-got-char":
17:32
<mpilgrim>
_(u"Expected closing tag. Unexpected character '%(data)s' found."),
17:32
<mpilgrim>
basically, search the file for occurrences of the "%" character
17:32
<mpilgrim>
some of them do provide quite a bit of value
17:32
<mpilgrim>
for example:
17:32
<mpilgrim>
"unexpected-end-tag":
17:32
<mpilgrim>
_(u"Unexpected end tag (%(name)s). Ignored."),
17:33
<mpilgrim>
i don't want to strip that information
17:33
<mpilgrim>
but string formatting in py3 is completely different
17:34
<jgraham>
Well it is not impossible to make some compatibility function for templating in either language
17:34
Philip`
wonders when an HTML WG telcon announcement email last got all the dates and times correct
17:35
<Philip`>
(or even just self-consistent)
17:37
<mpilgrim>
python3/html5lib/tokenizer.py is badly out of date
17:37
<mpilgrim>
but as far as i can tell, it only deals with unicode strings
17:39
<mpilgrim>
html5lib/tokenizer.py contains this line:
17:39
<mpilgrim>
char = eval("u'\\U%08x'" % charAsInt)
17:39
<mpilgrim>
which is quite funky and will fail in py3
17:39
<mpilgrim>
and will not be auto-converted to anything useful
17:40
<mpilgrim>
perhaps we can split that out into a separate function in a separate file
17:40
<mpilgrim>
because other than that, i believe tokenizer.py can be safely auto-converted
17:41
<jgraham>
That could probably be rewritten not to use string formatting
17:41
<mpilgrim>
or eval
17:41
<jgraham>
Yeah :)
17:42
jgraham
wonders who wrote that code in the first place
17:42
<jgraham>
I'm pretty sure it wasn't me
17:42
<mpilgrim>
html5parser.py is too badly out of date for me to tell whether auto-conversion would work
17:45
<mpilgrim>
actually, the old-style string formatting still works in python 3.1
17:45
<mpilgrim>
it's just deprecated
17:45
<mpilgrim>
so we can ignore the problems in constants.py, for now
17:45
<mpilgrim>
(though that eval still wouldn't work, since it uses the u"" form that no longer exists)
17:47
<mpilgrim>
html5lib/filters/* can all be safely auto-converted
17:47
<mpilgrim>
(lint.py uses %-style string formatting though)
17:52
<mpilgrim>
html5lib/serializer/htmlserializer.py contains a call to reduce()
17:52
<mpilgrim>
which should probably be rewritten with a for loop
17:53
<mpilgrim>
after which, html5lib/serializer/* could be safely auto-converted
17:54
<mpilgrim>
html5lib/treebuilders/__init__.py contains several import statements not on the top level
17:54
<mpilgrim>
those do not get auto-converted by 2to3
17:55
<Dashiva>
reduce is gone?
17:55
<Dashiva>
I'm so out of it
17:55
<mpilgrim>
Dashiva: http://diveintopython3.org/porting-code-to-python-3-with-2to3.html#reduce
17:55
<mpilgrim>
and those import statements are broken in py3
17:56
<mpilgrim>
they need to be rewritten using the new syntax for relative imports
17:56
<mpilgrim>
but since they're not at the top level of the module, 2to3 won't rewrite them
17:56
<mpilgrim>
not sure what to do about that, except fork it
17:57
<Philip`>
Write a custom preprocessor
17:58
<mpilgrim>
not sure it's worth it
17:59
<mpilgrim>
html5lib/treebuilders/_base.py can be safely auto-converted
18:02
<mpilgrim>
treebuilders/dom.py is out of date, but i believe it can be safely auto-converted
18:04
<mpilgrim>
treebuilders/etree.py is out of date, and it uses %-style string formatting, but i believe it can be safely auto-converted
18:06
<mpilgrim>
treebuilders/etree_lxml.py is out of date, and it uses %-style string formatting, but i believe it can be safely auto-converted
18:06
<mpilgrim>
i'm not entirely sure though
18:07
<mpilgrim>
it has some relative imports that 2to3 might choke on
18:07
<zcorpan_>
wouldn't the justgiving example in http://html5doctor.com/measure-up-with-the-meter-tag/ use <progress>?
18:08
<mpilgrim>
treebuilders/simpletree.py can be safely auto-converted
18:08
<TabAtkins>
zcorpan_: Yeah, that's a <progress>.
18:11
<zcorpan_>
what do people think should happen for data:text/xml,<!DOCTYPE x [<?xml-stylesheet href='data:text/css,x{background:papayawip}'?>]><x/>
18:11
<mpilgrim>
treebuilders/soup.py is out of date, and it uses %-style string formatting, but i believe i can be safely auto-converted
18:11
<zcorpan_>
xml-stylesheet 1st ed says it should be applied
18:12
<zcorpan_>
ie and webkit apply it, firefox and opera don't
18:12
<zcorpan_>
s/papayawip/papayawhip/
18:23
<TabAtkins>
zcorpan_: I don't see any particular reason to think it shouldn't be applied. Is the issue the nesting of data: urls, or just whether or not xml-stylesheet should accept data: urls?
18:23
<TabAtkins>
(In either case, I think it should be fine.)
18:23
<jgraham>
mpilgrim: To find out just what can't be auto-converted it might be better to start from the python 2 trunk -- except inputstream.py -- 2to3 everything else, see where things break, and try to fix up the python 2 source so they don't break anymore
18:23
<zcorpan_>
TabAtkins: no, the question is whether a PI in the internal subset should be applied as an xml-stylesheet PI for the document
18:24
<zcorpan_>
TabAtkins: or whether it should be ignored
18:30
<TabAtkins>
zcorpan_: Oh, okay. Sorry for the noise then. I have no opinion.
18:30
<zcorpan_>
ok
18:42
<Philip`>
zcorpan_: Why would it be ignored?
18:42
Philip`
doesn't see why it'd be different to e.g. inserting a <style> element via the internal subset
18:43
<zcorpan_>
Philip`: you can't have elements in the internal subset
18:44
<zcorpan_>
you can have an entity declaration in the internal subset and then an entity reference somewhere that expands to a <style> element, but that's not the same thing
18:45
<Philip`>
Oh
18:48
<zcorpan_>
PIs in the internal subset are not represented in the DOM at all, other than part of the internalSubset attribute which is just a DOMString
18:48
<zcorpan_>
(which is tentatively dropped in web dom core)
18:54
<mpilgrim>
jgraham: yeah, i was expecting more breakage than i found looking through diffs
19:03
<zcorpan_>
if someone can find any content that uses an xml-stylesheet PI in the internal (or external) subset, or can provide data about lack of such content, i would appreciate it
19:05
Philip`
can't, since he can barely even find page that use XHTML
19:05
<Philip`>
*pages
19:06
<borismus>
/j cmu
19:06
<erlehmann>
my page does. and occasionally breaks
19:08
<Philip`>
erlehmann: Your site is http://blog.dieweltistgarnichtso.net/ ?
19:08
<Philip`>
Even http://blog.dieweltistgarnichtso.net/?s=cheese breaks
19:08
<Philip`>
never mind http://blog.dieweltistgarnichtso.net/?s=%ef%bf%bf
19:08
<erlehmann>
every search term breaks it. havent updated the theme for a while.
19:09
<Philip`>
The second one breaks it more than the first, though :-p
19:10
<erlehmann>
damn wordpress
19:10
<Philip`>
erlehmann: Blame XML
19:10
<Philip`>
It makes these things far too complex
19:11
<erlehmann>
no i dont.
19:11
<erlehmann>
u just haven't got a good token-thingy
19:11
<erlehmann>
hopefully html5lib will change that
19:12
<erlehmann>
but the theme devs as well as wordpress devs CLAIMED it could do xhtml
19:12
Philip`
wouldn't mind draconian error handling of element syntax and nesting, because that's pretty trivial to get right; it's just the obscure little edge cases of invalid characters or forbidden sequences of characters or forbidden names or forbidden attribute values that are a huge pain to protect against
20:00
zcorpan_
doesn't see an internal subset in http://blog.dieweltistgarnichtso.net/
20:09
<Philip`>
zcorpan_: I think he just meant his page uses XHTML
20:57
<tantek>
greetings - anybody here know who has admin privs on http://wiki.whatwg.org/ ?
20:57
<tantek>
there's been a couple of spam pages added (see http://wiki.whatwg.org/wiki/Special:RecentChanges ) and I'd like to help out by deleting them and blocking the spammers
21:23
<jgraham>
tantek: I don't. I seem o recall that Lachy does but he is away. Hixie does too of course
21:23
<tantek>
thanks jgraham.
21:24
<tantek>
Hixie, Lachy, there's been a couple of spam pages added (see http://wiki.whatwg.org/wiki/Special:RecentChanges ) and I'd like to help out by volunteering to be a wiki admin to delete and block spammers.