skip to main |
skip to sidebar
by
Aaron Adler and Randal Davis Summary
Oltman attempted to use speech as a way to overcome the ambiguities of sketching. The system proposed was limited though. It was very domain specific, and communication only went one way. Adler and Davis hope to create a white board system that incorporates both sketch and speech to aid in early design work. The system should be able to engage the user in natural dialog.A user study was performed in an almost Wizard-of-Oz setup. The participant was given a variety of sketching goals to accomplish. The experimented had an identical tablet and engaged the user in dialog while the user was attempting to accomplish their goals. The participant was able to change the color of their strokes, and noted was that strokes fall into one of four categories: creation, modification, selection or writing. The authors were able to make three observations about the participants speech. First, they were difluent and often repeated words or phrases. Second, when prompted with a question from the experimenter they often responded with words that were used in the original question. Third, the speech utterance were related to what the user was drawing at the moment. Other observations were that the user would often list objects and then sketch them in that order, they wrote out words they used in their speech, and participants often paused their speeches so as to finish the drawing they were describing.When prompted with a question from the experimenter the user often gave much more elaborate answers than were necessary. Often they would spot errors or ambiguities when giving these responses. Participants also made comments not related to the sketch, but in relation to the domain. The paper goes on to give details about the connection between the time of the sketch and word/phrase groupings, noting that speech phrases preceded sketching.Discussion
There were a lot of observations presented that could be incorporated into a system. This paper seems to be the ground work for an implementation. I question whether having the experimenter in the room would have caused the participant to interact more? If the participants would be so giving if it was only them and a machine? I liked the use of color as a way to give context about the drawing. There could be additional layers to provide context that wouldn't necessarily need to show up on the sketch. Interested to see where they take this.Citation
A. Adler and R. Davis. Speech and sketching: An empirical study of multimodal interaction. In Fourth Eurographics Conference on Sketch Based Interfaces and Modeling, Riverside, California, August 2-3 2007.
by
Michael Oltmans and Randall DavisSummary
The paper combines speech recognition techniques and sketch recognition into a system titled ASSISTANCE. The system is able to recognize 2-D kinematic diagrams through the visuals provided through sketching and the context given by the user voicing additional information about what should occur in the sketch. Sketch recognition is in constant pursuit of supplementing the process that designers use in their early stages. This means low-level quick sketches, and also vocal discussions. The designer should be able to explain the intended actions by drawing additional context (arrows to indicate direction), pointing and giving spoken feedback. All of this information is used to build a model of the device, and the user could ask the system questions about the model. The paper tackles how to handle overlapping descriptions to infer meanings. ASSISTANCE takes three inputs, the structural model of the sketch generated by ASSIST, the phrases recognized by the speech recognition (ViaVoice) and the sketched arrows. ASSISTANCE must be able to unify multiple description instances, and resolve deictic references. The system must also understand casual links. Discussion
I like this paper because it presents the combination of multiple natural methods of interaction to create something that would often require a much more restrictive technique. It is a good starting point. I am unsure how the difference between pointing and drawing is done. Also, does one simply draw, switches modes, and then begins to explain (using the pointer to make references)? Or can the user give verbal interaction throughout the whole experience? What happens if a mistake has made? Can statements be made contradictory and only the latter statement is used?Citation
Oltmans, M. and Davis, R. 2001. Naturally conveyed explanations of device behavior. In Proceedings of the 2001 Workshop on Perceptive User interfaces (Orlando, Florida, November 15 - 16, 2001). PUI '01, vol. 15. ACM Press, New York, NY, 1-8. DOI= http://doi.acm.org/10.1145/971478.971498