00:01
<eseidel>
TabAtkins: well, I uploaded my first-pass tests to https://bugs.webkit.org/show_bug.cgi?id=45950. I've not yet run them. I'll work on more tests and implementation tomorrow. Still unclear how difficult this will be. The devil is in the details.
00:04
<TabAtkins>
kk, cool
08:33
<annevk>
ooh
08:34
<annevk>
was it webreq not pubreq?
08:34
<annevk>
yes it is :/
09:32
<annevk>
so it seems Gecko is the only browser to support http://en.wikipedia.org/wiki/Big5#Unicode-at-on which is cited as of limited success
09:32
<annevk>
taking IE's code table, then overriding the PUA characters with the HKSCS mapping, might be a somewhat sensible way forward
09:53
<hsivonen>
hmm. an HTML WG poll coming up. we haven't had one of those in a while
10:02
<zcorpan>
hsivonen: what's it about?
10:09
<hsivonen>
zcorpan: content model of <object>
10:10
<zcorpan>
woot
10:11
Ms2ger
sighs, ignores
10:12
<jgraham>
But isn't it obvious that a popular vote is the best way to decide technical questions? The only think we're doing wrong is not televising it. And getting Simon Cowell involved somehow.
10:13
<Ms2ger>
We've got a Simon, let him do it
10:16
<zcorpan>
Content model: a single sarcasm element.
10:17
<jgraham>
I'm not sure we would fit the role. zcorpan doesn't naturally inspire the urge to punch him in the face (at least as far as I can tell).
10:17
<jgraham>
s/we/he/
10:17
<zcorpan>
i guess i should work on that
10:20
<Ms2ger>
Yes! Then you can become a chair too!
10:45
<hsivonen>
TabAtkins: looks like your ISSUE-134 CP expired
10:45
<hsivonen>
this one http://lists.w3.org/Archives/Public/public-html/2011Jul/0235.html
10:45
<hsivonen>
per March 28 deadline
10:48
<annevk>
so IE PUA extensions to big5 are 0x81 / 0xA0 and 0xC6 / 0xC8
10:49
<hsivonen>
there's no agenda for the HTML WG f2f meeting, right?
10:53
<annevk>
haven't seen anything, but haven't paid to close attention either
10:53
<annevk>
s/to//
10:55
<zcorpan>
leif's CP says to allow nesting interactive content, but doesn't really explain why
10:56
<hsivonen>
annevk: ok. well, I think I won't be traveling to the meeting
10:57
<jgraham>
Was any justification given for the meeting for either HTML or WebApps?
10:57
<zcorpan>
"authors can just fill it with content without having to re-author the page first (that is: they don't need to make changes to the parent element in order to add the most optimal fallback content)." - not if you want e.g. a <form> in the fallback
10:58
<jgraham>
It kind of felt like "well everyone lives in the bay area anyway so we may as well get together"
10:58
<annevk>
so the incompatible overlap seems to be in 0xA3
10:58
<jgraham>
Even though that isn't true for the person organising it…
10:59
<annevk>
in big5 those are code points outside the PUA (at least in IE) and in hk they map to something else
10:59
<hsivonen>
jgraham: yeah, it felt like that
10:59
<annevk>
I was thinking of going
11:00
<annevk>
but if nobody is going to show up, hmm
11:00
<annevk>
I'm bad with jetlag these days
11:00
<hsivonen>
annevk: sounds bad considering that it seems that jetlag travel is your job
11:02
<jgraham>
hsivonen: One assumes that there will be some people that live within walking (en-us: driving) distance that will make it
11:03
<jgraham>
s/hsivonen/annevk/
11:05
<jgraham>
Not sure it is worth flying half way around the world just to see those people though
11:05
<annevk>
yeah, and I think I "have to" show up for WebApps
11:05
<jgraham>
Oh, really?
11:05
<annevk>
sort of told Art I would; he didn't think it would be useful without editors
11:06
<jgraham>
I see
11:06
<annevk>
on the upside it's over a month and I've no travel planned this one
11:06
<annevk>
I think that might be a first
11:11
<hsivonen>
you could have opted to show up at TPAC later this year and make jetlag Someone Else's Problem
11:11
<hsivonen>
or do what Hixie does: not showing up at meeting even if it's said that meeting without editors aren't useful
11:12
<annevk>
chrome and chrome-hk differ for 85-A0, A3, C6-C8 (with some compatible code points in between)
11:12
<hsivonen>
or was it axiomatic that there has to be a meeting?
11:13
<annevk>
not sure really; I kind of like visiting CA though and seeing everyone again
11:24
<annevk>
passion -> progress -> process -> !passion -> !progress
11:27
<zcorpan>
while (passion) make(progress);
11:28
<Ms2ger>
elihw; make(process);
11:29
<annevk>
what's elihw?
11:30
<MikeSmith>
hsivonen: WebApps WG is also meeting. People seemed to have found the f2f time useful at TPAC for discussing Web Components, among other things
11:33
<Ms2ger>
You mean XBL2 :)
11:46
<karlcow>
passion is a bad start… in a latin way
11:59
<MikeSmith>
anybody know if the minutes from today's IETF meetings are online yet?
12:05
<MikeSmith>
finds http://tools.ietf.org/wg/httpbis/minutes
12:06
<Ms2ger>
"mark observed that the problem of on-the-wire bits being turned into uris is now a problem for someone else"
12:06
<Ms2ger>
"larry masinter challenges the assertion that iri are working on this"
12:06
<Ms2ger>
Go Larry!
12:07
<MikeSmith>
:)
12:07
<MikeSmith>
http://tools.ietf.org/wg/hybi/minutes
12:07
<zcorpan>
annevk: maybe build a list of URLs looking for the relevant byte sequences for pages that have a big5 label, and check them manually
12:07
<MikeSmith>
"Action Item: Chairs to coordinate with W3C WebApps WG on API impact of MUX"
12:07
<Ms2ger>
Nooooooooooo
12:09
<hsivonen>
what's MUX?
12:09
<annevk>
zcorpan: if we can assume they're all authored for a single UA that might work...
12:09
<MikeSmith>
multiplexing?
12:09
<karlcow>
sounds an electronic music soundtrack
12:09
<karlcow>
sounds like*
12:09
<hsivonen>
MikeSmith: did hybi add multiplexing to Web Sockets?
12:10
<karlcow>
https://en.wikipedia.org/wiki/Multiplexer
12:10
<annevk>
hsivonen: not yet, I think, but I'm not following hybi
12:10
<karlcow>
http://www.w3.org/Protocols/MUX/
12:10
<hsivonen>
don't they respect the layered architecture where the nature of the IP layer takes care of multiplexing? :-)
12:11
<Ms2ger>
Zing.
12:11
<karlcow>
ahah ♥ http://www.w3.org/Protocols/MUX/WD-mux-980708.html
12:11
<hsivonen>
karlcow: thanks
12:11
<karlcow>
SMUX Protocol Specification
12:11
<karlcow>
W3C Working Draft 08-July-1998
12:11
<annevk>
we should have a bot that flags users when they start talking about layered architectures
12:11
MikeSmith
finds http://tools.ietf.org/html/draft-tamplin-hybi-google-mux-03
12:11
<MikeSmith>
expired though
12:12
<MikeSmith>
hmm yeah
12:12
<MikeSmith>
http://tools.ietf.org/wg/hybi/agenda?item=agenda-83-hybi.html
12:12
<Ms2ger>
MikeSmith, clearly the W3C document isn't expired, so it must be more authoritative
12:13
<MikeSmith>
za-zing
12:13
<karlcow>
but but but I love Millefeuilles http://farm3.staticflickr.com/2454/3883365345_ea76895ae0_m.jpg
12:13
<annevk>
zcorpan: who handles our tools for that internally?
12:13
<MikeSmith>
http://tools.ietf.org/agenda/83/slides/slides-83-hybi-0.pdf
12:13
<annevk>
found the interface...
12:14
<Ms2ger>
annevk, ... you've got internal tools for tracking academics?
12:16
<hsivonen>
huh? Is PDF legal at the IETF now? I insist on ASCII .txt slides!
12:16
<MikeSmith>
:)
12:16
<MikeSmith>
http://tools.ietf.org/html/draft-montenegro-httpbis-speed-mobility-01 references the websockets multiplexing I.D.
12:16
<annevk>
when Mike asked about minutes I tried looking something up and found that HTML + style sheets would be converted to plain text, but HTML without style sheets was fine
12:17
<annevk>
clearly the idea that style sheets are optional is not a concept they grasp
12:17
<annevk>
I think that was here: http://www.ietf.org/wg/meeting-materials.html
12:17
<annevk>
Philip`: do you still have your largish dataset?
12:18
<MikeSmith>
wtf
12:18
<jgraham>
annevk: I might worry that these datasets will be quite western-biased
12:18
<MikeSmith>
"Minutes may be submitted in plain text (Unix/Mac/Dos), in simple HTML with NO style sheets, PDF, and PPT. Minutes submitted with style sheets WILL BE converted to plain text. Therefore, if you are concerned about preserving the formatting of your minutes, then please submit them in simple HTML."
12:18
<annevk>
MikeSmith: yeah
12:18
<hsivonen>
what bundled app can I use on Windows 7 to test if the virtual machine deals with microphones reasonably?
12:18
<zcorpan>
annevk: why would you need to assume that?
12:18
<annevk>
MikeSmith: they're hiliarious
12:19
<zcorpan>
annevk: what tools?
12:19
<hsivonen>
hmm. maybe I should just install Skype in the VM and see if I can make a test call
12:19
<jgraham>
MikeSmith: That's … well it's funny for one thing. But how wrongheaded do you have to be to accept PPT but not HTML + CSS?
12:19
<annevk>
zcorpan: guess we can have people that read those variants of Chinese have a look
12:20
<MikeSmith>
I think these are the guidelines from bizarro superman planet
12:20
<hsivonen>
MikeSmith: CSS not ok but PPT ok. That's sad.
12:20
<Ms2ger>
IETF. That's sad.
12:21
<MikeSmith>
Bizzarro Superman say PPT GOOD but HTML BAD for minute of Internet standard meeting
12:21
<annevk>
zcorpan: was thinking of MAMA
12:21
<zcorpan>
annevk: ah. i think mama isn't really maintained anymore. :(
12:21
<Ms2ger>
twss
12:23
<annevk>
jgraham: the bias matters less as we'll only look at the big5 subset
12:23
<MikeSmith>
anyway, Yebisu time
12:24
<karlcow>
echo "IETF" | sed -e 's/IE/W/'
12:24
<jgraham>
annevk: Sure, you might just find that there isn't all that much data
12:25
<karlcow>
annevk: each time I asked about MAMA, dad told us, she was in limbo.
12:30
<Philip`>
annevk: Not in a conveniently accessible form (it's in some backup archive on some disk or other)
12:30
<Ms2ger>
Philip`, how about canvas tests?
12:32
<annevk>
Philip`: how big is the dataset?
12:34
<Ms2ger>
Sorry annevk, I'm afraid I scared him off
12:34
<Philip`>
annevk: It's http://webcrawl.s3.amazonaws.com/web200904.gz so only 2.6GB
12:35
<Philip`>
(though I didn't use that file directly, I parsed the HTTP headers and split it in multiple files for easier processing)
12:35
<zcorpan>
annevk: grep -aPibo "(content(-type\s*:|=\s*[\"']?)\s*text/html\s*;|<meta\s)\s*charset\s*=\s*[\"']?(big5|cn-big5|csbig5|x-x-big5|big5-hkscs)" web200904
12:35
<jgraham>
annevk: I thought zcorpan had been using the dotbot data
12:35
<jgraham>
Oh dammit
12:35
<Ms2ger>
Are you parsing HTML with regices, sir?
12:36
<jgraham>
Ms2ger: In other news, yay!
12:36
<zcorpan>
http://dotnetdotcom.org/
12:36
<annevk>
zcorpan: hmm, why don't you give me the results? :)
12:36
<zcorpan>
Ms2ger: yeah, i also do paris hilton (but don't tell my wife)
12:36
<Ms2ger>
jgraham, mm?
12:37
<zcorpan>
annevk: it gives byte offsets so you need to download the data to interpret the result anyway :-P
12:37
<jgraham>
Ms2ger: You Resolved:FIXED the import testharness.js bug
12:37
<Ms2ger>
Ah, yes
12:37
<zcorpan>
maybe i should do like Philip` and split the data into multiple files
12:38
<Ms2ger>
Now we will run 15 getElementsByClassName tests we didn't before!
12:38
<Philip`>
(I did the splitting mainly so I could easily run multiple threads that each pick a big chunk of pages and process them all)
12:39
<jgraham>
Ms2ger: Baby steps :)
12:39
<Ms2ger>
Indeed
12:39
<annevk>
zcorpan: so then write some other scripts that uses those offsets to read into the data and analyze for bytes?
12:39
<annevk>
hmm
12:40
<Ms2ger>
Now, if only Aryeh's tests didn't take so long that debug builds seem unable to handle them...
12:40
<zcorpan>
dunno, maybe the grep is the wrong starting point
12:40
<annevk>
sounds like pain and also like it will take a long time
12:41
<annevk>
my laptop is already slow with 1MB of JSON :)
12:41
<Philip`>
grep seems like a bad idea when you're trying to process whole pages, not just individual lines, and when you want to understand HTTP headers properly instead of getting confused by pages that happen to quote HTTP headers in their body text
12:42
<MikeSmith>
perl -pi -e and undef $/
12:43
<MikeSmith>
though that still doesn't hep you for docs that quite HTTP headers in their body
12:44
<MikeSmith>
or i-less even
12:44
<Philip`>
You mean read the entire ~14GB uncompressed input file into a string in RAM?
12:44
<MikeSmith>
if you only grepping
12:44
<MikeSmith>
yeah
12:44
<Philip`>
Even I wouldn't use Perl for that
12:44
<MikeSmith>
cause that's what real men do
12:44
<annevk>
haha, need some more RAM first
12:45
<MikeSmith>
when I need more RAM I just beat somebody and takes theirs
12:46
<MikeSmith>
btw, I added lightning bolts to http://platform.html5.org/
12:48
zcorpan
runs grep -aEizHZ "(content(-type[[:space:]]*:|=[[:space:]]*["']?)[[:space:]]*text/html[[:space:]]*;|<metas)[[:space:]]*charset[[:space:]]*=[[:space:]]*["']?(big5|cn-big5|csbig5|x-x-big5|big5-hkscs)" web200904 > big5.txt
12:48
<zcorpan>
it forgets about the URLs, but maybe that's ok
12:49
<annevk>
I guess we can first see if there's any data at all
12:49
<zcorpan>
that uses big5 encoding declaration?
12:49
<zcorpan>
seems to be plenty
12:51
<MikeSmith>
https://twitter.com/#!/mikebelshe/status/185650746623135745
12:53
<annevk>
is that concept "None of us is as dumb as all of us." explained somewhere?
12:53
<annevk>
apart from the motivational poster
12:53
<hsivonen>
I see Hixie made the blockquote in http://hsivonen.iki.fi/producing-xml/ non-conforming :-(
12:54
<MikeSmith>
annevk: it started from that motivational poster I think
12:54
<Ms2ger>
A long time ago, no?
12:55
<MikeSmith>
annevk: though it reminds me of "We have met the enemy and he is us."
12:55
<MikeSmith>
http://en.wikipedia.org/wiki/Pogo_(comic_strip)#.22We_have_met_the_enemy_and_he_is_us..22
13:00
<karlcow>
the French quote is the dumb thing is "L'intelligence de la foule" ~ aka Intelligence of the crowd… meant to qualify things like Heysel Stadium disaster
13:13
<zcorpan>
annevk: http://simon.html5.org/dump/big5.txt (29MB) - each response body is prefixed by "web200904\0"
13:14
<annevk>
hsivonen: there's already https://developer.mozilla.org/en/Document_Object_Model_(DOM)/window.navigator.language
13:14
<annevk>
zcorpan: sweet, so that's a subset basically?
13:15
<zcorpan>
annevk: yep
13:15
<annevk>
very interesting
13:15
<zcorpan>
shouldn't take too long to run a python script over it on a laptop
13:16
<annevk>
I'm going to buy some lunch first I think, but that should indeed be fine
13:16
<annevk>
thanks
13:16
<zcorpan>
np
13:18
<zcorpan>
it may contain false positives, like pages that have two declarations or just has content-type... as text in the page, but i guess that's pretty rare
13:24
MikeSmith
wonders if jzaefferer is around
13:30
<zcorpan>
annevk: hmm, it seems i may have got the regexp wrong. it doesn't seem to include content-type: matches
13:30
<zcorpan>
annevk: only content="..." matches
13:34
<zcorpan>
the problematic part seems to be the colon
13:35
<SHAGGSTaRR>
colons are always fulla shit zcorpan
13:56
<annevk>
zcorpan: guess I'll try to make something that works with this first
13:59
zcorpan
now runs grep -aEizHZ "(Content-Type[[:space:]]*:[[:space:]]*text/html[[:space:]]*;[[:space:]]*charset[[:space:]]*=[[:space:]]*["']?|[[:space:]]content[[:space:]]*=[[:space:]]*["']?[[:space:]]*text/html[[:space:]]*;[[:space:]]*charset[[:space:]]*=[[:space:]]*["']?|<meta[[:space:]]+charset[[:space:]]*=[[:space:]]*["']?)(big5|cn-big5|csbig5|x-x-big5|big5-hkscs)" web200904 > big5.txt
14:07
<jzaefferer>
hey MikeSmith, I'm here
14:13
<annevk>
so far I find trouble
14:14
<[tm]>
jzaefferer: away from my PC now but just wanted to day i will have something fire you this weekend or Monday
14:15
<[tm]>
something for you
14:15
<Ms2ger>
You will have something fire him? Tough man
14:15
<[tm]>
Heh
14:16
<[tm]>
you prefer to script this on node i guess?
14:16
<annevk>
what's the problem with
14:16
<annevk>
i = 0
14:16
<annevk>
prev_byte = None
14:16
<annevk>
for b in bytes_read:
14:16
<annevk>
if prev_byte != None:
14:16
<annevk>
if ord(prev_byte) > 0xA0:
14:16
<annevk>
i+=1
14:16
<annevk>
#if ord(prev_byte) ==0xA3 and ord(b) > 0x3F and ord(b) < 0xFF:
14:16
<annevk>
# i+=1
14:16
<annevk>
prev_byte = None
14:16
<annevk>
continue
14:16
<annevk>
if b < 0x80:
14:16
<annevk>
continue
14:16
<annevk>
prev_byte = b
14:17
<annevk>
apart from that > 0xA0 gives less results than < 0xA1!!!!
14:17
<scott_gonzalez>
[tm]: Yeah, our build script is written in node.
14:17
<[tm]>
ok
14:22
<hsivonen>
annevk: right. I'm skeptical about the usefulness of exposing a language list in a new API
14:22
<[tm]>
so what we need to do us use whatever nodes http client is to send th contents of each file as part of a post request to port 8888
14:23
<[tm]>
loop through arts from the command line
14:24
<[tm]>
and do an async read
14:24
<[tm]>
since it's node
14:25
<[tm]>
anyway i can write it but i guess you guys could too
14:25
<[tm]>
better than me
14:26
<[tm]>
so maybe i will just get the jars together and give you the details about how to start up the service and what exactly to send in the post
14:27
<hsivonen>
I wonder why Opera Mini on tablets doesn't say "Tablet" in the UA string like Opera Mobile on tablets
14:28
<annevk>
missed an ord()
14:28
<annevk>
still there's a lot of code points apparently that are not interoperable at all
14:33
<hsivonen>
hmm. Hixie's response to the responsive images threads wasn't particularly apt
14:33
<jgraham>
He wasn't responsive enough?
14:33
<hsivonen>
Hixie: you really don't see a use case for adapting what number of pixels you send depending on the recipient?
14:34
<hsivonen>
Hixie: sending the "same" photo sampled differently
14:34
<hsivonen>
Hixie: e.g. Flickr varies the bitmaps for the "same" photo it sends depending on context
14:35
<hsivonen>
Hixie: but they have to do imperatively instead of having a declarative way of doing it
14:35
<hsivonen>
(in their slideshows, for example)
14:36
<karlcow>
[10:28] <hsivonen> I wonder why Opera Mini on tablets doesn't say "Tablet" in the UA string like Opera Mobile on tablets
14:36
<karlcow>
I wonder if it's a bug or not.
14:36
<karlcow>
It's always a lot of discussions
14:37
<zcorpan>
annevk: ok i've uploaded a new big5. although it looks like it's the same size, so either it was right the first time around, or it's still wrong, or it's now right but it didn't make any difference...
14:37
<zcorpan>
gotta go
14:37
<hsivonen>
oh well. there's the poll: https://www.w3.org/2002/09/wbs/40318/issue-158-objection-poll/
14:37
<hsivonen>
from September 2002
14:39
<matjas>
fun fact: IE < 8 doesn’t recognize `\:` as a CSS escape sequence for `:`, so you have to use `\3a ` instead.
14:40
<karlcow>
hsivonen: why do you think it's from September 2002
14:41
<hsivonen>
karlcow: is your IRC client in Quebec?
14:41
<annevk>
thanks simon
14:41
<hsivonen>
karlcow: /2002/09/
14:41
<jgraham>
karlcow: Because normal people (and Google) don't assume that URIs are opaque
14:41
<karlcow>
why do you think this is a date
14:41
<annevk>
i somehow thought extracting useful data was going to be easier, but it's still quite a lot of data
14:41
<hsivonen>
karlcow: it looks like one
14:41
<karlcow>
hsivonen: not to me ;)
14:41
<jgraham>
karlcow: Are you disputing that it is, in fact, a date?
14:41
<hsivonen>
karlcow: you've been on the W3C staff
14:42
<jgraham>
(just not a date that corresponds to anything relevant)
14:42
<annevk>
most of the W3C staff realizes that it's silly though
14:42
<annevk>
that's why we have /html/ now
14:42
<annevk>
and /ns/
14:42
<karlcow>
jgraham: yes ☺
14:43
<annevk>
claiming URLs are opague when they are in front of people all the time seems extremely counter productive
14:43
<karlcow>
annevk: tell that to mobile users, twitter and other shorteners
14:43
<annevk>
especially as they are, unlike say barcodes, very readable and to some extent memorable
14:43
<hsivonen>
people who believe URLs are opaque should use IP addresses instead of hostnames in their URLs
14:44
<hsivonen>
IP address, a slash and an UUID
14:44
<annevk>
karlcow: they didn't exist when this argument was made up
14:44
<karlcow>
and? :D
14:44
<annevk>
karlcow: and twitter uses it because <a ping> got canned
14:45
<annevk>
that some URLs are less readable than others, does not make them opague
14:45
<hsivonen>
annevk: good point. let's blame URL shorteners on people who objected to <a ping>
14:45
<karlcow>
huh? ☺ I don't understand the drift here
14:47
<karlcow>
domain names are definitely not opaque for most people so far, except when using QR code. (another layer), but the path is completely lost for most of the people I know who are not in the computing industry.
14:47
<karlcow>
Geek Magnifying glass.
14:48
<hsivonen>
karlcow: http://picturesofpeoplescanningqrcodes.tumblr.com/
14:48
<karlcow>
what I see is a lot of people using search engines with keywords for finding stuff they already know
14:49
<jgraham>
Sure, I do that. I also expect that if I find http://some.blog.com/2011/11/21/some_entry.html it was published in November 2011
14:49
<karlcow>
hsivonen: yes QR codes have a very very very limited area of real usage. ☺ the marketing saga here is insane.
14:49
<karlcow>
jgraham: disqualified. Geek! ;)
14:49
<annevk>
per this dataset the C6-C8 range does not matter much
14:49
<jgraham>
karlcow: Possibly, but you don't actually have any evidence
14:50
<annevk>
there is one instance of such a lead byte and it's inside a comment
14:50
<jgraham>
For your assertions
14:50
<karlcow>
jgraham: this goes both ways
14:50
<jgraham>
Sure.
14:51
<annevk>
advertising frequently uses the path
14:51
<annevk>
example.com/boldnewplan
14:51
<hsivonen>
facebook.com/aolkeyword
14:57
<annevk>
Content-Language: big5
14:58
<annevk>
this data contains fun stuff
15:00
<annevk>
also apparently 74944 bytes following a lead byte that are less than 0x40 (and thereby invalid)
15:03
<annevk>
although that's on 1761026 lead bytes
15:09
<michel_v>
what karlcow meant about dates in URLs is that the 2002/09 may only be the date that the resource was created on the server
15:09
<michel_v>
it might have changed a lot since then, with updated data. but its URL didn't because it's cool like that
15:12
<annevk>
https://www.w3.org/2012/03/12-ab-minutes (W3C Member-only)
15:13
<jgraham>
Why would anyone put a date in a URL for a resource that will change? (In this case the "resource" from 2002 is presumably the wbs system itself, not any of the surveys. Which is kind of like putting the date that Wordpress was first deployed in the url of all wordpress blogs)
15:14
<scott_gonzalez>
[tm]: Sorry, wasn't looking at this channel. Yeah, if you can just send us the details about how to start the service and make the request, we can write the node code.
15:14
<michel_v>
because it's a naming convention that just happens to be different from the mainstream conventions
15:15
<Philip`>
jgraham: Some people are terrible at coming up with names, and would rather only have to compete with other people's names from the same month rather than trying to be unique across all of W3C history
15:15
<hsivonen>
annevk: ooh. a date in a W3C URL!
15:15
<karlcow>
example.org/637GFDRkdsat5-suyig.html
15:15
<michel_v>
jgraham: (besides, on blogs the resource behind a "dated" URL does change over time while the date does not. that's comments)
15:16
<annevk>
hsivonen: don't tell karlcow
15:16
<Philip`>
(The technology society's fixation on TLAs doesn't help with the name collision problem)
15:16
<michel_v>
(or other ways to influence the resource that are not the author's actions)
15:16
<jgraham>
michel_v: I'm not sure how that is relevant. The date the article was first published is useful irrespective of later comments.
15:17
<karlcow>
"Ceci n'est pas une date." — Magritte
15:17
<jgraham>
The date that the underlying software was first deployed, not so much
15:17
<jgraham>
karlcow: Somewhat lacks to double-entendre that I am told is implied in the original
15:18
<jgraham>
s/to/the/
15:19
<karlcow>
I will use another forbidden word here ;) the calendar is a nice way to make layers of unique identifier because of the arrow of time.
15:19
<karlcow>
identifiers
15:26
<annevk>
hmm
15:26
<annevk>
no lead bytes in the 0xFx range
15:27
<annevk>
that is somewhat surprising
15:27
<hsivonen>
annevk: have you tested Python and Java big5 decoding?
15:28
<annevk>
hsivonen: nah, just browsers
15:29
<annevk>
not sure what we'd gain from those libraries, presumably they're even less compatible with what's out there
15:38
<annevk>
so in this dataset, there's 602746 lead bytes less than 0xA1 (undefined territory for big5) and 20825 0xA3 lead bytes followed by reserved territory
15:38
<annevk>
there's also 74374 invalid trail bytes
15:39
<annevk>
and a total of 1113300 lead bytes over 0xA1 (you'll have to subtract the 0xA3 bytes from that)
15:40
<annevk>
anyone ideas on how to analyze this?
15:54
<annevk>
I guess I could make character buckets
15:54
<annevk>
but actually, you probably need sequences to make sense of the Chinese
15:54
<annevk>
foolip: you around?
16:05
<sedovsek>
Hey! Any ideas on how to find webpages that are using multiple column layout? Either column-count: <integer> or column-width: <size>;
16:05
<sedovsek>
Manythanks!
16:05
<annevk>
there's one item here that has iso-8859-1 in HTTP and then in the HTML it has
16:05
<annevk>
<meta http-equiv="content-type" content="text/html; charset=euc-kr"/>
16:05
<annevk>
<meta http-equiv="content-type" content="text/html; charset=EUC-JP"/>
16:05
<annevk>
<meta http-equiv="content-type" content="text/html; charset=big5-hkscs"/>
16:05
<annevk>
all three
16:05
<sedovsek>
Sorry, didn't know I was bursting into the conversation.
16:05
<annevk>
HTML paradise
16:05
<annevk>
sedovsek: it's more of a monologue, no worries
16:06
<sedovsek>
Oh. :)
16:11
<annevk>
sedovsek: dataset we're using is from 2009; not sure if there's anything more recent
16:11
<annevk>
sedovsek: oh, and I think URLs are stripped
16:12
<annevk>
heh, the above was the only big5-hkscs declaration
16:12
<annevk>
so big5-hkscs is just not present
16:13
<sedovsek>
香港增補字符集 :P
16:15
<sedovsek>
I did some presentation about multiple columns, regions, exclusions & stuff like that ... and would like to show some real examples of, at least multi columns.
16:15
<sedovsek>
But found only 6 so far.
16:15
<sedovsek>
Which is quite sad because it degrades nicely.
16:15
<annevk>
hmm, only site I knew that was using it appears no longer to use it ( http://robert.ocallahan.org/ )
16:16
<smaug____>
MikeSmith's spec list uses it
16:16
<annevk>
ah yeah, sedovsek http://platform.html5.org/
16:16
<sedovsek>
Many thanks!
16:17
<sedovsek>
http://veerle.duoh.com/about - bio section here as well.
16:18
<sedovsek>
And wikipedia for references.
16:29
<karlcow>
http://my.opera.com/hallvors/blog/2012/03/29/slashdot-runs-out-of-slashes
16:32
<sedovsek>
Not sure if this is the right channel to ask this question, but ...
16:32
<sedovsek>
how come column-span has values either none or auto.
16:33
<sedovsek>
I would expect it could be a number as well?
16:33
<sedovsek>
For instance ... I want this h1/h2 element to span only across two columns.
16:33
<sedovsek>
(in case of 3 or more column layout)
16:40
<annevk>
too complex for the first version or something like that
16:40
<TabAtkins>
Yeah, values other than "1" and "all" were allowed at first, but it complicates layout something fierce.
16:41
<TabAtkins>
(Fucking floats.)
16:41
<sedovsek>
It crossed my mind, yea.
16:41
<sedovsek>
Especially with fluid column widths, flexible column counts, etc.
16:42
<TabAtkins>
And things where the spanner appears halfway in the multicol element, rather than at the beginning.
16:42
<TabAtkins>
With the current one, if a spanner appears you can just chop the multicol element into two multicol segments with the spanner between them, much simpler.
16:44
<sedovsek>
True, but that might require additional markup?
16:44
<sedovsek>
But it's also a fallback solution (for browsers that does not support column-span).
16:44
<TabAtkins>
No, I mean that the *layout engine in the browser* can do that.
16:45
<TabAtkins>
It's relatively simple to do that, from a layout perspective, than to deal with an element intruding across some of the columns.
16:45
<sedovsek>
Oh.
16:47
<sedovsek>
Another question if may I ...
16:48
<sedovsek>
just wrote docs for column span yesterday
16:48
<sedovsek>
https://developer.mozilla.org/en/CSS/column-span
16:48
<sedovsek>
Wasn't sure about column-span support in different browser.
16:48
<sedovsek>
Opera supports it from 11.1+?
16:48
<TabAtkins>
No clue about support. I know Opera probably supported it first, given that it's Hakon's spec.
16:49
<sedovsek>
Thank you.
16:50
<smaug____>
I thought it was in Gecko first, but could be wrong
16:50
<smaug____>
ah, column-span
16:51
<smaug____>
roc was hacking columns quite a bit at some point
16:57
<annevk>
charset=x-MS950-HKSCS o_O
16:58
<smaug____>
that looks interesting
16:58
<annevk>
ooh, there's a bunch more HKSCS data if you look for BIG5-HKSCS
16:58
<annevk>
my hex editor does case-sensitive searching, which is not that surprising I suppose
16:58
<annevk>
maybe it's more surprising that people use uppercase for these things
16:59
<annevk>
oh, Big5-HKSCS and MS950 also exist
16:59
<annevk>
charset=null :)
17:01
<annevk>
the ms950 is a Java invention apparently that somehow leaked
17:20
<dglazkov>
good morning, Whatwg!
17:22
<smaug____>
um, yeah, this is the way to write specs. let's not define what the API do but "Alternatively, you could say that the current webkit implementation is the reference." :/
17:22
<smaug____>
Web Audio is so under-specified
17:25
<annevk>
big5 is too
17:25
<TabAtkins>
smaug____: Argh, that's quite bad.
17:26
<smaug____>
TabAtkins: the spec defines all sorts of audionodes which modify the data somehow, but it is not defined how
17:26
<smaug____>
how can anyone write tests for that
17:27
<smaug____>
how can any web dev rely on the API
17:27
<annevk>
so people do actually have stuff like:<!-- meta http-equiv="content-type" content="text/html; charset=big5" -->
17:27
<jgraham>
smaug____: That's not a literal quote is it?!
17:28
<smaug____>
jgraham: the part inside "" is
17:28
<annevk>
heh, and in this entire dataset charset is never followed by a space
17:28
<annevk>
maybe that's a bug in the regexp
17:29
<annevk>
smaug____: euh wut, pointer?
17:29
<annevk>
sounds like vp8
17:30
<smaug____>
annevk: it is here http://lists.w3.org/Archives/Public/public-audio/2012JanMar/0543.html
17:31
<smaug____>
I know Raymond had other suggestions too
17:31
<smaug____>
but is felt very strange that one can even write such idea that let's use some implementation as reference and not really specify stuff
17:32
<jgraham>
Oh, it's not a quote from the spec. Well that's a little better
17:32
<smaug____>
jgraham: oh, that would be strange language in a spec
17:33
<smaug____>
is anyone from Opera in the Audio WG?
17:33
<jgraham>
Well the spec does mention webkit
17:33
<jgraham>
I... don't think so. foolip perhaps?
17:33
<annevk>
don't think we're in the WG, but reverse engineering the market leader is not something we're interested in
17:33
<annevk>
that's why we have standards...
17:36
<annevk>
<cfprocessingdirective pageencoding="Big5-HKSCS"> sweet
17:38
<annevk>
so out of 964, 14 pages declare big5-hkscs
17:39
<annevk>
and one of those does so via ColdFusion :p
17:43
<jgraham>
Oh foolip is in the group
17:43
<jgraham>
But he is stretched pretty thin :(
18:10
<annevk>
does html5lib do character encoding detection correctly?
18:10
<annevk>
hmm, but I can't actually use html5lib
18:11
<annevk>
I need a byte-level tokenizer
18:11
<annevk>
man :(
18:11
<annevk>
I guess it should be easy enough to adapt the html5lib tokenizer to work on bytes
18:12
<annevk>
but it's getting pretty convoluted just to analyze some data
18:12
<Philip`>
Can you just decode as iso-8859-1 chars then tokenise?
18:13
<annevk>
probably
18:13
<annevk>
the other problem is actually analyzing the data
18:14
<annevk>
there's about 630000 code points that will map to PUA in IE but are potentially better decoded as HKSCS
18:15
<annevk>
I guess if foolip is around I can try feeding him the first 100 or so and see whether it makes sense to continue and do something more elaborate
18:16
<annevk>
and I should probably generate the ~1000 files so I can inspect them each individually and make sure they are indeed encoded as big5/big5-hkscs and not utf-8 or some such
18:35
<bga>
is it true that you want to replace http?
18:37
<izhak>
no that's not true
18:37
<TabAtkins>
I want to replace http.
18:37
<annevk>
some people do
18:37
<TabAtkins>
With telepathy.
18:37
<izhak>
whatwg claims nothing about http
18:37
<TabAtkins>
So probably not a practical desire.
18:38
<TabAtkins>
I'm very confused as to where this meme came from, bga.
18:38
<annevk>
prolly the HTTP 2.0 / SPDY stuff
18:38
<bga>
but spdy, webrtc, http://blogs.msdn.com/b/interoperability/archive/2012/03/26/speed-and-mobility-an-approach-for-http-2-0-to-make-mobile-apps-and-the-web-faster.aspx
18:38
<TabAtkins>
spdy has absolutely nothing to do with whatwg.
18:38
<TabAtkins>
webrtc isn't whatwg anymore either.
19:44
<annevk>
no
19:44
<annevk>
we need to define the concept it seems
19:45
<annevk>
oh
19:53
<zewt>
there's no single place in every url-fetching api that's correct for that feature
19:53
<rubys>
Is Hixie around?
19:54
<zewt>
also it's much more than a "nit", heh
19:58
<rubys>
When Hixie gets back, I'd like to talk more about http://krijnhoetmer.nl/irc-logs/whatwg/20120328#l-1010 ; I talked to people I expected to balk at it, and to my surprise, it might be doable. I just have a few questions, and perhaps we can make it happen as soon as Monday.
19:59
<arun__>
zewt, you're right, it's more than a nit. Do you think it's worth defining it for Blob URIs in FIle API? I'm inclined to think so.
19:59
<otherarun>
(note my new nick ^)
20:00
<zewt>
i think the "automatically-released blob urls" feature is worth doing it, yes (though as I mentioned later in the thread, I think the "release at a later point" instead of "on first use" is a better way of doing it)
20:00
<zewt>
(of course, it's not me that has to do the work, so take my "worth it" with whatever weight you like :)
20:01
<otherarun>
I think it's worth defining 'dereferencing.' That shouldn't be bandied about loosely.
20:01
<zewt>
i'd definitely say that if we can't or won't define it, then the auto-releasing thing needs to be dropped entirely, since it's doomed to not being interoperable
20:02
<otherarun>
zest, I agree. I just worry that it is a harder problem than it looks like.
20:02
<zewt>
well, it looks like each spec that takes a URL and initiates a fetch will need to invoke "dereference" explicitly somewhere
20:02
<otherarun>
gah, ^ zewt I mean
20:03
<otherarun>
I wonder what's rigorous enough here.
20:03
<zewt>
for example, "let urlRef be the dereferenced value of url", where urlRef is logically either 1: the underlying blob data or 2: the url itself (if it's just a regular URL)
20:03
<zewt>
then, urlRef is what's passed to the fetch algo
20:03
<zewt>
i'm sure there are plenty of hard parts of actually doing that, since so many places fetch
20:03
<otherarun>
Right. Where here 'fetch' is the dereferencing steps.
20:04
<zewt>
no, "fetch" is the "fetch a resource" algorithm
20:04
<zewt>
http://www.whatwg.org/specs/web-apps/current-work/multipage/fetching-resources.html#fetch
20:04
<zewt>
afk a few
20:07
<jgraham>
Hmm, nobody told rubys that Hixie is away
20:07
<jgraham>
Maybe he will read the logs
20:07
<otherarun>
zewt, I think we can put dereference in terms of fetch. More thought needed on this.
20:07
<otherarun>
zewt, but the call for doing it more rigorously or bailing is a good one.
20:11
<zewt>
otherarun: this must be done synchronously, before the surrounding JS call returns; fetch is often done asynchronously, or in a queued task
20:11
<otherarun>
hmmm
20:11
<zewt>
seems overly restrictive to try to require every usage of fetch be initiated synchronously
20:12
<zewt>
(we should have this discussion when anne or hixie are around)
20:13
<zewt>
basically, though, "dereference" can basically transform the URL into something that (at the spec level) can be used like the URL as far as fetch is concerned, but actually holds a reference to the underlying blob data
20:14
<zewt>
there could be a better way, of course
20:15
<otherarun>
zewt, that's probably what'll end up being the case. I'm not sure exactly how much normalization the Blob URI needs to undergo.
20:15
<otherarun>
(note that the entire protocol is defined like GET is)
20:15
<zewt>
what do you mean "normalization" exactly?
20:16
<otherarun>
Oh, I mean what you mean when you say "basically transform the URL into something that (at the spec level) can be used lie the URL as far as fetch is concerned"
20:17
<otherarun>
^ lie = like above
20:17
<zewt>
in JSish pseudocode, it'd be something like var someUrl; urlOrBlob = dereference(urlOrBlob); fetch(urlOrBlob);
20:17
<MikeSmith>
to his surprise, heh
20:17
<otherarun>
MikeSmith, oh hai
20:18
<zewt>
where dereference would be eg. function dereference(url) { if is_a_blob_url(url) return theBlob; else return url; }
20:18
<otherarun>
zewt, aha!
20:18
<zewt>
(except it would dereference to the *underlying* data of the blob, not the actual blob, so it's unaffected by blob.close() and transfers)
20:18
<otherarun>
zewt, what happens on unsuccessful deref?
20:18
<zewt>
(we may also need some clear way of referring to "the underlying data of a blob")
20:18
<zewt>
hmm. leave the URL as-is, so you get an error at fetch time?
20:19
<zewt>
(the same error you get now, whatever that is)
20:19
<otherarun>
Yeah, ok.
20:19
<otherarun>
(Well you get a 500)
20:19
<otherarun>
(cribbing from HTTP parlance)
20:19
<zewt>
(404 seems to make more sense, but it's not terribly important)
20:19
<zewt>
(since it should never really happen in non-buggy code anyway)
20:20
<MikeSmith>
try to constrain discussion to some very narrow outcome and the express "surprise" when someone says "hey let's maybe not constrain the discussion to your artificially constrained outcome"
20:20
<MikeSmith>
surprise, surprise
20:20
<zewt>
what battle are you fighting? heh
20:21
<otherarun>
Some concern existed about information reveal about underlying filesystem, so chose to make invalid retrievals return 500s
20:21
<zewt>
can't imagine how that could happen, but it's not important enough to worry about now
20:21
<otherarun>
+1
20:22
<zewt>
hmm
20:22
<zewt>
well
20:23
<zewt>
if a spec is *already* initiating the fetch synchronously, this is probably unneeded
20:23
<zewt>
i don't know if anyone does that
20:23
<otherarun>
In our case, fetches are synchronous
20:23
<zewt>
anyone = specs
20:25
<zewt>
for example, XHR send() in async mode goes asynchronous before initiating the fetch
20:26
<zewt>
(as opposed to starting a fetch with the fetch synchronous flag unset, so fetch step 4 would do it)
20:26
<otherarun>
I think that's part of the problem with use of terms like 'dereference.' It traipses over sync/async considerations.
20:27
<zewt>
(actually, it does both, depending on the code path ... don't recall why)
20:28
otherarun
= afk
20:38
<rubys>
jgraham: do you know when Hixie is expected back?
20:38
<TabAtkins>
saturday or sunday
20:38
<rubys>
ok, will try back then. Thanks!
20:51
<sedovsek>
Here's the short list of websites that are using multiple column layout:
20:51
<sedovsek>
http://galjot.si/multiple-columns#current_use
21:00
<zewt>
those layouts are usually really terrible ... that one at the top is utterly unreadable
21:04
<sedovsek>
I would like to comment (either agree or disagree) on that, but since i've only found 7 examples ... :)
21:05
<sedovsek>
Can't really tell whether they're terrible or not.
21:08
<ojan>
TabAtkins, Hixie: what's the rationale behind limiting the number of on* event handlers we add? Why not just support on* events for all events?
21:08
<TabAtkins>
I don't know. I'd like to add them for everything.
21:08
<TabAtkins>
Hixie's the one who wants to limit them.
21:09
<ojan>
The only think I can think of is the fear that we'll conflict with some existing on* attribute that the existing pages use...so then we'd have to name the event something less optimal for back-compat.
21:09
<ojan>
that doesn't seem like a great reason though
21:10
<isherman>
TabAtkins: any news on https://bugs.webkit.org/show_bug.cgi?id=66032 ?
21:10
<TabAtkins>
Not from me. ;_;
21:11
<isherman>
k, I'll try to track down Tantek sometime…
21:18
<Velmont>
Is ;_; crying? Like ^_^ unhappy: ·_· and with tears?
21:19
<TabAtkins>
Yes.
21:19
<Velmont>
Nice.
21:19
<TabAtkins>
;_; and T_T are both crying.
21:25
<Velmont>
sedovsek: Ofc there's much more usage than that. I've used it for several sites a long time. -- Biggest problem ofc is that they continue too long. So will always have to "guarantee" that the text isn't higher than a screenful.
21:26
<sedovsek>
Velmont: nice point regarding text height.
21:26
<isherman>
tantek: I think I'm supposed to ping you about https://bugs.webkit.org/show_bug.cgi?id=66032 — any chance of this getting some attention soon from the CSS working group? (sorry if that's not quite the right name for the group, I'm not entirely up to speed here)
21:26
<TabAtkins>
Yeah, until multicol has (1) the ability to wrap to a new "row" when your height is constrainted and you've filled up your width, and (2) the ability to *not* columnate when the content is below a certain length, it's not really usable for arbitrary web paes.
21:26
<Velmont>
TabAtkins: Yep. So not for template driven stuff.
21:27
<TabAtkins>
Yeah.
21:27
<TabAtkins>
I use it on an personal web page to flow a whole bunch of recipe names across the page.
21:27
<tantek>
Hi isherman
21:27
<Velmont>
I've recently seen the light though, and will do more custom stuff going forwards. It's so freeing to actually be able to customize more stuff by writing special css and markups for different pages.
21:27
<tantek>
checking your www-style message now
21:27
<TabAtkins>
But it's okay for that to get taller than the page, because they're not meant to be read straight through.
21:27
<isherman>
tantek: thanks :)
21:28
<Velmont>
TabAtkins: Yep, - like wikipedia's references. That is a good usage.
21:28
<sedovsek>
TabAtkins, Velmont: I've used it here: http://galjot.si/talks/future-layouts/#slide30
21:28
<tantek>
isherman I believe there is some work being done on this
21:28
<TabAtkins>
Yeah, exactly like Wikipedia.
21:28
<sedovsek>
But that's because I need not care, or at least not so much, about degradation.
21:28
<sedovsek>
Wikipedia is great example, yes.
21:29
<isherman>
tantek: any way for me to stay in the loop on that work?
21:29
<tantek>
isherman ok found where it's being tracked for CSS4-UI: http://wiki.csswg.org/spec/css4-ui#more-selectors
21:29
<tantek>
:autofill is what you're asking about right?
21:30
<isherman>
tantek: yep, that's exactly it
21:30
<isherman>
tantek: would it be appropriate to add support for this selector to WebKit, or is it too early for now?
21:30
<tantek>
I presume you mean prefixed, like :-webkit-autofill
21:32
<isherman>
tantek: sure, prefixed if that's still the recommendation (I seem to recall having seen a long thread on WhatWG wondering about the value vs. cost of prefixes)
21:32
<tantek>
WhatWG has no authority on prefixing CSS features, the CSSWG does ;)
21:33
<tantek>
and yes, there's been lots of permathreads on several mailing lists about prefixing, few of them actually concluding with anything actionable or reasonably summarized.
21:35
<isherman>
But in the case of this selector, adding the prefixed ":-webkit-autofill" selector sounds appropriate? If so, I'll go ahead with that :)
21:35
<Velmont>
sedovsek: Found a page I randomly remember from a few years ago that I thought used it, but I see now it's just placed divs. Although it could've used it.
21:35
<tantek>
yes - implementations are always encouraged to innovate in ways that make sense to them. it helps more sensibly shape the standards :)
21:36
<tantek>
isherman, there's a couple of paths to advance this pseudo-class, I don't have a preference, but you may have an opinion. either we can develop it in CSS4-UI, or, because it is a selector, it can also go into the next version of Selectors.
21:36
<sedovsek>
If I compare this one http://platform.html5.org/ in multi column layout or IE9 (fallback) ...
21:36
<sedovsek>
I kinda like it better with IE9 ... it feels like it's easier to find links you are looking for.
21:37
<isherman>
tantek: Hmm, I'm not sure what the difference between those two paths is — are they both targeted at CSS4 (or whatever the next version of CSS will be named)?
21:39
<tantek>
one difference is that there is a FPWD of selectors4: http://www.w3.org/TR/selectors4/ whereas I have yet to write-up even an editor's draft of CSS4-UI (I'm still working on wrapping up CSS3-UI, now that its second LC has closed).
21:39
<tantek>
so selectors4 is further along from that perspective, and thus putting the feature there may get it a) into an editor's draft sooner, b) into a public WD sooner, and perhaps even LC/CR sooner.
21:40
<isherman>
tantek: well, I guess a faster path is better from my perspective :)
21:41
<tantek>
but these things of course depend on individual editors, their time/interest etc. however since fantasai is editing Selectors4 I expect it to proceed reasonably swiftly (unless she takes on more work that slows her down) and likely faster than I can get CSS4-UI to the same draft state.
21:42
<isherman>
that makes sense
21:43
<isherman>
I think either way sounds fine to me — I'll leave it to your discretion, since you're much more familiar with this process
21:43
<tantek>
I'd also encourage lurking in irc://irc.w3.org:6665/css
21:43
<isherman>
but I'll go ahead and push forward on getting the prefixed version into WebKit
21:44
<tantek>
makes sense
21:44
<tantek>
and will help add weight/incentive/priority to get it spec'd
21:46
<isherman>
perfect — thanks again for the advice :)
21:47
<tantek>
isherman - no problem at all - thanks for the heads up.
21:48
<tantek>
isherman - I've also filed a bugzilla bug for the :-moz-autofill equivalent in Gecko: https://bugzilla.mozilla.org/show_bug.cgi?id=740979
22:31
<smaug____>
no aklein