VSS 2025 Abstracts


Jump to:
 
Bai, D., Weisman, K., Schille-Hudson, E., Ihm, E., Taves, A., Luhrmann, T. M., & Scholl, B. J. (2025). Chasing the supernatural: The perception of animacy from motion is related to spiritual experiences. Poster presented at the annual meeting of the Vision Sciences Society, 5/20/25, St. Pete Beach, FL.  
When simple geometric shapes move about a display, what do we see? Beyond the obvious lower-level properties (e.g. direction, velocity), we may also spontaneously and even irresistibly perceive seemingly higher-level properties in certain displays -- such as one object *chasing* another. Here we asked whether such percepts might relate to a rather different form of phenomenology: the experience of supernatural or spiritual presences. If some people move through their daily lives perceiving higher degrees of agency, could this also manifest in more reports of sensing the presence of gods or other spiritual beings? Whereas some past work has theorized about such connections, here we isolate visual processing (vs. higher-level interpretation) by exploring correlations with objective visuomotor detection and performance. Subjects viewed displays filled with multiple identical moving discs. One (the 'sheep') was controlled by their own mouse movements, while another (the 'wolf') pursued the sheep. Subjects could only avoid being 'caught' by identifying the wolf's motion amidst the many distractors. Afterwards, subjects also completed two prominent and nuanced measures of spiritual experience (the 'Spiritual Events scale', and the 'Inventory of Non-Ordinary Experiences'). The results were striking: subjects who reported spiritual or supernatural experiences (via both measures) were worse at detecting chasing (and thus avoiding the wolf) -- and this difference in psychophysical functions could not be explained by either demographic variables (e.g. age, education) or more generic measures of religious belief or practice. We explain this by appeal to a type of 'hyperactive agency detection': people who perceive agency even in 'distractors' (and are therefore worse at detecting chasing) are more likely to report spiritual experiences. This work thus highlights a striking connection between a form of basic visual detection and some of the most powerful experiences in many people's lives.
 
Colombatto, C., McCrackin, S., Scholl, B. J., & Ristic, J. (2025). Eyes vs. attentiveness: Pupil dilation widens the perceived cone of direct gaze. Talk given at the annual meeting of the Vision Sciences Society, 5/17/25, St. Pete Beach, FL.  
One of the most important social signals we perceive is the direction in which another person looks, especially when they are looking at us. In fact, humans perceive a range of eye-gaze deviations as directed towards (vs. away from) us, a range known as the Cone of Direct Gaze (CoDG). Does the CoDG reflect perceived eye direction per se, or might it indirectly reflect the degree to which we see another person *attending* to us? We explored this by asking whether the CoDG is affected by another salient property of others' eyes often linked to attention -- how dilated the pupils are. We reasoned that perceived pupil dilation (vs. constriction) may lead to increased perception of the gazer attending to the observer, resulting in a wider CoDG. We tested this idea in two preregistered experiments (N=150 each). Observers viewed faces with constricted, normal, or dilated pupils, embedded in eyes looking with varying eccentricities to the left, right, or directly at the observer. The manipulation of pupil dilation was entirely task-irrelevant, as observers' task was simply to report the gaze direction that they saw. In an initial experiment, faces with dilated pupils were more likely to be perceived as gazing directly toward the observers, compared to faces with either normal or constricted pupils. When the faces were inverted in another experiment, however, this effect vanished -- suggesting that the impact of pupil dilation on the CoDG is not just a function of the brute physical differences of the pupils themselves (which of course still exist even when inverted). We interpret this effect -- that pupil dilation widens the CoDG -- in terms of enhanced percepts of attentiveness. Judgments of gazing direction are thus influenced not just by physical properties of the eyes, but also by higher-level percepts of the cognitive states behind the eyes.
 
Erdogan, M., Nobre, A. C., & Scholl, B. J. (2025). Can you attend broadly in space while attending narrowly in time?: On the generality of attentional breadth. Talk given at the annual meeting of the Vision Sciences Society, 5/20/25, St. Pete Beach, FL.  
A salient aspect of spatial attention is its variable *breadth*: sometimes we select narrowly (e.g. when hunting for lost keys in a cluttered drawer), while other times we select broadly (e.g. when viewing the overall configuration of players on a soccer field). And an analogous dynamic applies across time: sometimes we attend to relatively high-frequency changes (e.g. when listening to fast jazz) while other times we focus on events changing at a slower pace (e.g. waves crashing ashore). How are these different forms of attentional breadth related? For example, can you attend broadly in space while simultaneously attending narrowly in time? One might have no effect on the other. Or attending broadly in one domain might facilitate attending broadly in the other. Or broad vs. narrow attention might draw in part on different resources, such that attending broadly in one domain is easier when attending narrowly in the other. Participants were presented with four types of stimuli at once: (a) visual probes flashed in a relatively narrow ring around fixation (spatially narrow), (b) visual probes flashed in a wider ring relatively far from fixation (spatially broad), (c) repeated-tone probes in a high-frequency auditory stream (temporally narrow), or (d) repeated-tone probes in a lower-frequency stream (temporally broad). Across trials, participants were instructed to attend broadly vs. narrowly in space, and (independently) broadly vs. narrowly in time. As expected, attending to one visual ring impaired performance in the other -- and ditto for the two auditory streams. Critically, there were also cross-dimension interactions: for example, participants were better at focusing spatial attention *broadly* (detecting probes in the wider ring) when their temporal attention was focused *narrowly* (detecting probes in the higher-frequency stream). The interplay between spatial and temporal attention may thus depend on its relative breadth in each domain.
 
Ji, H., & Scholl, B. J. (2025). 'Visual verbs' drive adaptive predictions: Perception of dynamic event types spontaneously changes visual working memory encoding. Talk given at the annual meeting of the Vision Sciences Society, 5/18/25, St. Pete Beach, FL.  
We see the world not only in terms of specific features (such as the color or shape of a ball), but also in terms of a foundational set of abstract/categorical 'event types' (such as a ball bouncing vs. rolling). Recent work has demonstrated that such categorical perception occurs spontaneously during passive viewing of visual scenes, even when verbal encoding is discouraged or disrupted: observers are better able to detect changes across different event types, even when the magnitudes of within-type changes (e.g. across two different animations of bouncing) are objectively greater. Why might this occur? Here we explored the possibility that such spontaneous categorical encoding is adaptive, insofar as it enables differential predictions about likely future states, and so changes what is encoded into memory. This was inspired by the idea that the purpose of perception is not only to characterize the present ("What's out there?") but also to predict the future ("What's about to happen?"). We studied this in a single-trial memory task, e.g. when contrasting bouncing vs. rolling animations: observers viewed a single animation of a ball moving, and then simply reported its final position (after the video had ended and the display had disappeared). This placement was systematically biased by the underlying event type: rolling balls tended to be localized as further back in their actual trajectories horizontally (but not vertically), compared to bouncing balls -- presumably because a bouncing ball can only move forward in the coming moments, while a rolling ball could roll backwards down a ramp. And careful controls showed that this depended on the event-type itself, rather than any lower-level properties (such as the details of the trajectories). This shows how representations of 'visual verbs' might drive adaptive predictions about how a dynamic world is likely to unfold.
 
Jones, H., Bai, D., Scholl, B. J., & Awh, E. (2025). Electroencephalogram decoding of the attentional selection and tracking of featureless objects. Poster presented at the annual meeting of the Vision Sciences Society, 5/19/25, St. Pete Beach, FL.  
Recent work leveraging multivariate decoding of EEG data has identified a signal that scales with the number of items in working memory (WM). This signal appears to be content-independent, generalizing across distinct visual features, and between single-feature and multi-feature items. One explanation for this content independence is that this signal reflects an abstract indexing process that binds items to their context in space and time for maintenance and access. To explore this possibility, we examined whether a similar load signal exists for "featureless objects", which have no enduring properties from moment to moment. On each trial, participants viewed a dense grid of crosses of random orientations. A moving object was implemented by having a cross change from one random orientation to another, with these changes propagating through space and time. These transients yield a persisting trackable object, even though (a) there is no surface feature that is constant from one frame to the next, and (b) it is impossible even in principle to identify an object in any static frame. In the actual experiment, participants viewed a set of such moving featureless objects at once, and were cued to track 1 or 2 of them. After tracking the cued item(s), there was a brief delay, and then participants had to discriminate between the true final location of an object, and a nearby alternative location. EEG decoding found a signal that scaled with the number of featureless objects and generalized between all 3 phases of the trial: cueing, tracking, and the pre-test delay. In next steps, this neural signature will be compared directly to the previously identified WM load signal. Generalization of these load signals across task contexts would provide support for the theory that spatiotemporal indexing plays a role in WM maintenance.
 
Verosky, N., & Scholl, B. J. (2025). An event sequence to remember: Abstract temporal structure influences memorability. Poster presented at the annual meeting of the Vision Sciences Society, 5/19/25, St. Pete Beach, FL.  
Much work in recent years has explored how some stimuli are intrinsically more memorable than others, often due to distinctive sensory or semantic features. But what determines memorability beyond such features, e.g. in sequences of musical tones? Here we explored the possibility that memorability also depends on abstract temporal structure, in both vision and audition. Participants were presented with short temporal sequences of three items varying along a scalar dimension--either three successive tones of different pitches (in auditory experiments) or three successive circles of different sizes (in visual experiments). Items were sampled from five fixed points along the relevant scalar dimension, creating a combinatorial space of 60 possible sequences in each modality, with a one-to-one mapping between analogous auditory and visual sequences. (Examples of possible sequences would thus be 1-2-3, 1-3-2, 1-3-5, and 5-2-4--with the items 1-5 mapping onto either the pitches of tones or the sizes of circles.) To comprehensively characterize memorability, we presented each participant with sequences that were randomly sampled from the full combinatorial space. During stimulus presentation, participants made an orthogonal judgment with no mention of memory. Subsequently, they completed a surprise recognition task. Though all sequences were constructed from the same small library of meaningless items, some were nevertheless reliably more memorable--and some particular sequences were exceptionally memorable across both modalities. For example, among all sequences in the combinatorial space, those that stood out as especially memorable were monotonically increasing sequences consisting either of three consecutive items (1-2-3, 2-3-4, 3-4-5) or of the lowest-magnitude plus the two consecutive highest-magnitude items (1-4-5). Intriguingly, this same pattern emerged independently for both vision and audition. This work thus adds a new dimension to the study of memorability: beyond distinctive sensory and semantic properties, what we incidentally remember is also shaped by abstract temporal structure.
 
Walter-Terrill, R., & Scholl, B. J. (2025). Spatial affordances trigger spontaneous visual perspective-taking even in the absence of other agents. Poster presented at the annual meeting of the Vision Sciences Society, 5/18/25, St. Pete Beach, FL.  
A central goal of vision is to recover information about local environments that is not tied to a single perspective ("What does it look like from here?"), but is generalizable ("What's out there?"). Most work on such themes involves either deliberate imagery ("What would it look like from over there?") or social cognition -- as when taking another agent's perspective ("What would it look like from her shoes?"). Here, in contrast, we show that spontaneous visual perspective-taking occurs even in the absence of other agents -- triggered just by the spatial affordances of the environment itself. Observers viewed two buttons on a table and pressed keys with their right or left hands to indicate when those buttons haphazardly changed to particular colors. The two buttons were vertically aligned from the observers' viewpoint (one closer, one further away), such that adopting a perspective from the left or right side rendered responses spatially congruent or incongruent (as in the Simon effect). Critically, this different perspective was afforded by an environmental regularity in the scene that influenced whether the buttons were reachable. For example: (1) One side of the table had a chair, with the other side flush against a wall. (2) One side of the table had a normally-oriented chair, while the other had a backwards-facing chair. (3) One side of the table had no obstruction, while the other had a translucent screen. Or (4) one side of the table was on solid ground, while the other stood over a sheer cliff face. Each case yielded robust spontaneous visual perspective-taking: despite the task-irrelevance of these manipulations, observers responded faster when the afforded perspectives yielded congruent spatial button/response mappings. These results show how spatial affordances spontaneously promote a type of generalizability during scene perception, and how visual perspective-taking does not require other agents.
 
Wong, K. W., Shah, A. D., & Scholl, B. J. (2025). Seeing from the ground up: Spontaneous perception of 'causal history' due to intuitive physics. Poster presented at the annual meeting of the Vision Sciences Society, 5/20/25, St. Pete Beach, FL.  
We typically think of visual perception as providing us with representations of our present local environments. But vision may also sometimes represent the causal *past*, extracting how those environments got to be that way -- as when a shape with a jagged 'bite' is represented as the full (unbitten) shape to which an event (biting) occurred. Here we suggest that this perception of 'causal history' is more prevalent than previously suspected, due to intuitive physics. In a stack of blocks (or books, or dishes), for example, gravity entails that the bottom object was placed before higher objects. Here we show that such 'historical' relationships are spontaneously extracted during passive viewing, and influence perception in surprising ways. Observers viewed a table on which a stack of two blocks appeared, (1) all at once, (2) with the bottom block appearing first, or (3) with the top appearing first -- and they simply reported on each trial whether the blocks appeared simultaneously or sequentially. We reasoned that possibility (3) might be less naturally perceived, since it violates the causal history mandated by the underlying intuitive physics. And indeed: when the bottom block appeared first, observers reliably perceived this sequential presentation; but when the top block appeared first, they were more likely to mistakenly perceive that blocks had appeared simultaneously -- as if the actual temporal offset and the gravity-inspired prior effectively cancelled out. In fact, observers were more accurate for towers built 'from the ground-up' than for actual simultaneous presentations (which were often misperceived as sequential). And these effects seemed specific to gravity-based intuitive physics, since they disappeared when the same stimuli appeared to be lying flat on the table. These results collectively show how visual processing extracts causal history as a result of intuitive physics, and how such representations influence the perception of temporal order.