What
Reported upstream as google/google-java-format#413 (open since 2019-11-05): a 6 kB file of commented-out code ran out of memory with a 12 GB heap, and its output would have been about 3 GB. It reproduces here on main (53ac7d8). The shape is block comments that end and start on the same line, as in commented-out code that had javadoc in it:
/*
class X {
*//** javadoc *//*
public void foo(Bar bar) {}
*//** javadoc *//*
public void foo(Bar bar) {}
*//** javadoc *//*
public void foo(Bar bar) {}
*//** javadoc *//*
public void foo(Bar bar) {}
}*/
The continuation lines of these comments come out indented by 17, 50, 116 and 248 columns: each comment on the line starts twice as far right as the one before it. The file size grows the same way:
| comments |
input |
output |
| 6 |
317 B |
4.3 kB |
| 10 |
517 B |
68 kB |
| 16 |
817 B |
4.3 MB |
Why it happens
Comment.computeBreaks moves the column past a comment by adding the length of the comment's last line:
int firstLineLength = text.length() - Iterators.getLast(Newlines.lineOffsetIterator(text));
return state.withColumn(state.column() + firstLineLength)
For a one-line comment that is right. For a multi-line one, JavaCommentsHelper.rewrite has just re-indented the continuation lines to start at state.column(), so the last line already contains the column, and it is counted twice. The next comment on that line is re-indented from the doubled column, and so on. Despite its name, firstLineLength holds the length of the last line.
google-java-format has the same code in Doc.Tok.computeBreaks.
Fix
After a comment that spans lines, the column is the length of its last line; after a one-line comment it stays column + length.
The output changes only where something follows a multi-line comment on that comment's last line, so a corpus run should show next to nothing.
Done when
- a golden of the file above indents each comment's continuation lines to the column where that comment starts, and a second run changes nothing;
- a test formats twenty such comments into a small output (the old code writes about 70 MB);
- a corpus run says how many files change.
What
Reported upstream as google/google-java-format#413 (open since 2019-11-05): a 6 kB file of commented-out code ran out of memory with a 12 GB heap, and its output would have been about 3 GB. It reproduces here on
main(53ac7d8). The shape is block comments that end and start on the same line, as in commented-out code that had javadoc in it:The continuation lines of these comments come out indented by 17, 50, 116 and 248 columns: each comment on the line starts twice as far right as the one before it. The file size grows the same way:
Why it happens
Comment.computeBreaksmoves the column past a comment by adding the length of the comment's last line:For a one-line comment that is right. For a multi-line one,
JavaCommentsHelper.rewritehas just re-indented the continuation lines to start atstate.column(), so the last line already contains the column, and it is counted twice. The next comment on that line is re-indented from the doubled column, and so on. Despite its name,firstLineLengthholds the length of the last line.google-java-format has the same code in
Doc.Tok.computeBreaks.Fix
After a comment that spans lines, the column is the length of its last line; after a one-line comment it stays
column + length.The output changes only where something follows a multi-line comment on that comment's last line, so a corpus run should show next to nothing.
Done when