A language model reliably catches structural problems: unclear claims, missing context, buried arguments, inconsistent framing. It cannot verify facts, judge whether something is true to your experience, or reliably assess unusual work, which it tends to read as unclear. Treat its notes as a careful first reader, not a verdict.
It can tell whether your writing is clear, consistent and structurally sound, which is part of good but not the whole of it. It cannot judge whether the work is true, necessary or worth anyone attention, and those are usually what determine whether something matters.
Because unfamiliar work often reads to a model as unclear rather than as new. It evaluates against patterns in existing text, so a genuine departure from those patterns can be penalised for exactly the quality that makes it interesting.
Only where the note identifies a real problem a reader would also hit. Rewriting to raise a score produces work that satisfies the rubric and nobody else, and the score rising can disguise the work getting duller.
No. It has no way to verify claims and will generally treat what you assert as given. Accuracy remains entirely your responsibility, and polished feedback on an inaccurate piece can be misleading precisely because it sounds thorough.